Dockerized Laya prediction
service: loads one or more checkpoints behind a router, then serves typed
decisions (choice / score / noul) over HTTP, auto-routed by language or pinned
with model. Also supports TypeSafe API compatibility via /v1/systemone.
Built and published for linux/amd64 and linux/arm64.
- 🔀 Multi-Checkpoint Routing: bundles
english,multilingualandtyped-decisions, auto-selected by language or pinned per request. - 🎯 Typed Decisions:
choice/score/noulquestions with calibrated probabilities, confidence and action probability. - 🤝 TypeSafe API Compatibility: drop-in
/v1/systemoneendpoint compatible with TypeSafe request/response schemas. - 🔐 Timing-Safe Auth: API keys (
X-API-Key/Authorization: Bearer) and HTTP Basic, compared in constant time. - 📑 Interactive OpenAPI Docs: Swagger UI (
/docs), ReDoc (/redoc) and the raw schema at/openapi.json. - 🧰 Built-in Presets: ready-made question sets (
triage,email,guard,moderation,router). - 📦 Bulk Inference:
/predict/bulkover many states, with per-state questions/model and isolated errors. - 🔎 Detection & Email Helpers:
/detect(script/language) and/email/state(clean + structure an email). - 🧩 Flexible State: string, JSON object, or conversation turns; criteria values may be any JSON (dicts/lists/numbers are rendered as compact JSON).
- 🛡️ Non-Root: runs as unprivileged
appuser(uid10001). - 📦 Multi-Architecture: supports both
linux/amd64andlinux/arm64. - 🩺 Healthcheck: dedicated
/healthzendpoint and containerHEALTHCHECK. - 🪶 CPU-Only Torch: uses the PyTorch CPU wheel index (no CUDA), roughly
0.35sper predict on CPU. - 🎮 Optional CUDA Variant: build the same Dockerfile with a CUDA
TORCH_INDEX(cu126by default) to install the CUDA torch wheel for GPU inference. - 🧠 Memory-Aware Loading: skips random weight initialization before loading checkpoints and trims the heap afterwards, avoiding the OOM spike on large encoders.
Pull and run the published image:
docker pull ghcr.io/chneau/laya
docker run -d -p 8000:8000 -e API_KEYS=key1 -v hf-cache:/data/hf ghcr.io/chneau/layaMulti-architecture images (linux/amd64 and linux/arm64) are published to GitHub Container Registry under several tags:
| Image Tag | Preloaded Models | Image Size | Description |
|---|---|---|---|
ghcr.io/chneau/laya:latest (or 0.6.1) |
None (Dynamic) | ~300 MB | Slim / Default: Small image size. Downloads model on first run into /data/hf. |
ghcr.io/chneau/laya:english |
english |
~1.3 GB | Instant Startup (English): Pre-baked English checkpoint, offline-ready. |
ghcr.io/chneau/laya:multilingual |
multilingual |
~1.8 GB | Instant Startup (Multilingual): Pre-baked multilingual checkpoint. |
ghcr.io/chneau/laya:typed-decisions |
typed-decisions |
~1.3 GB | Instant Startup (Typed Decisions): Pre-baked typed decisions checkpoint. |
ghcr.io/chneau/laya:all |
All 3 models | ~3.5 GB | Full Bundle: All checkpoints pre-baked for zero-latency multi-model routing. |
The same variants are also published with CUDA torch for linux/amd64 only
(PyTorch's CUDA wheels are not available for arm64), as :cuda, :english-cuda,
:multilingual-cuda, :typed-decisions-cuda and :all-cuda. They need the NVIDIA
container runtime at run time. See CUDA (GPU) Variant.
docker run -d -p 8000:8000 -e API_KEYS=key1 ghcr.io/chneau/laya:englishcurl http://localhost:8000/healthz- Swagger UI: http://localhost:8000/docs
- ReDoc: http://localhost:8000/redoc
- Raw OpenAPI Specification: http://localhost:8000/openapi.json
The examples/ folder contains ready-to-run HTTP requests and
a small Python client that classifies documents into categories. It only calls the
HTTP API and does not load checkpoints, so no local model setup is required.
| Method | Path | Auth | Description |
|---|---|---|---|
GET |
/healthz |
no | Liveness + resident models |
GET |
/models |
yes* | Available and loaded checkpoints |
GET |
/presets |
yes* | Built-in question sets |
POST |
/detect |
yes* | Script/language detection (what routing uses) |
POST |
/email/state |
yes* | Clean + structure an email as a state |
POST |
/predict |
yes* | Typed questions over one state |
POST |
/predict/bulk |
yes* | Same questions over many states |
POST |
/v1/systemone |
yes* | SystemOne / TypeSafe compatible prediction |
* Enforced only when API_KEYS and/or BASIC_AUTH is set. Any of these works:
-H 'X-API-Key: key1'
-H 'Authorization: Bearer key1'
-u admin:secret # HTTP Basiccurl -X POST localhost:8000/predict -H 'X-API-Key: key1' -H 'Content-Type: application/json' -d '{
"state": "I was billed twice. Please refund the duplicate today.",
"questions": {
"department": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "refunds", "technical": "bugs", "sales": "purchases"}},
"urgency": {"type": "score", "instructions": "How urgent?", "criteria": ["not urgent", "soon", "critical"]},
"refund": {"type": "noul", "instructions": "Does the customer ask for money back?"}
}
}'{
"model": "laya-rl-agent",
"answers": {
"department": {"type": "choice", "choice": "billing", "probabilities": {"billing": 0.96, "technical": 0.02, "sales": 0.02}, "confidence": 0.82},
"urgency": {"type": "score", "score": 1.36, "legend": {"0": "not urgent", "1": "soon", "2": "critical"}, "confidence": 0.09},
"refund": {"type": "noul", "noul": 0.82, "confidence": 0.82}
},
"usage": {"input_tokens": 132, "output_tokens": 0},
"routing": {"model": "english", "reason": "English Latin text"}
}state accepts a string, JSON object, or list of conversation turns. Criteria
values may be strings or any JSON value (dicts/lists/numbers are rendered as
compact JSON), and noul accepts optional {"true": ..., "false": ...} text.
Routing — omit model to auto-select by language (see /detect), or pin
"model": "english" | "multilingual" | "typed-decisions". The response includes
routing with the chosen checkpoint and reason.
Presets — skip questions and pass "preset": "triage" (one of triage,
email, guard, moderation, router); list them at GET /presets.
Bulk — use states with shared questions/preset/model, or items to
override questions and model per state. Returns {"count": N, "results": [...]}
with per-state errors isolated as {"ok": false, "error": "..."}.
curl -X POST localhost:8000/v1/systemone -H 'Authorization: Bearer key1' -H 'Content-Type: application/json' -d '{
"state": "I was billed twice. Please refund the duplicate today.",
"model": "laya-english",
"questions": {
"department": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "refunds", "technical": "bugs", "sales": "purchases"}},
"urgency": {"type": "score", "instructions": "How urgent?", "criteria": ["not urgent", "soon", "critical"]},
"refund": {"type": "noul", "instructions": "Does the customer ask for money back?"}
}
}'{
"model": "laya-english",
"answers": {
"department": {"type": "choice", "choice": "billing", "probabilities": {"billing": 0.96, "technical": 0.02, "sales": 0.02}, "confidence": 0.82},
"urgency": {"type": "score", "score": 1.36, "legend": {"0": "not urgent", "1": "soon", "2": "critical"}, "confidence": 0.09},
"refund": {"type": "noul", "noul": 0.82}
},
"usage": {"input_tokens": 132, "output_tokens": 0}
}Supported model names for TypeSafe requests:
laya-english, laya-multilingual, laya-typed-decisions.
| Variable | Default | Description |
|---|---|---|
API_KEYS |
(empty) | Comma-separated keys; empty disables API-key auth. |
BASIC_AUTH |
(empty) | Comma-separated user:password pairs. |
MAX_BULK_ITEMS |
(empty / unlimited) | Optional limit on states per /predict/bulk (unlimited by default). |
PORT |
8000 |
HTTP port (host and container). |
MODELS |
english |
Checkpoints to preload: english, multilingual, typed-decisions. |
MODEL_ID |
convaiinnovations/laya |
Optional repo override (mirror/local path). |
MODEL_SUBFOLDER |
(empty) | Subfolder for the English checkpoint. |
DEVICE |
cpu |
cpu or cuda. |
HF_HOME |
/data/hf |
HF cache (persisted via volume). |
HF_TOKEN |
(empty) | Optional Hugging Face token for authenticated checkpoint downloads. |
Preload MODELS=english,multilingual to route languages without reloading. On
startup the app also loads the repository .env (without overriding existing process
variables), so make run picks up local settings automatically.
services:
laya:
image: ghcr.io/chneau/laya
ports:
- "8000:8000"
environment:
- API_KEYS=replace_with_your_strong_api_key
- MODELS=english,multilingual
volumes:
- hf-cache:/data/hf
restart: unless-stopped
volumes:
hf-cache:cp .env.example .env # set API_KEYS
make up # docker compose up -d --buildThe default image is CPU-only. For GPU inference, build the same Dockerfile with
TORCH_INDEX pointing at the CUDA wheel index (cu126, torch 2.14.0) and
TORCH_DISABLE_NATIVE_JIT=1 so eager CUDA ops use prebuilt kernels instead of Triton
JIT builds. It needs the NVIDIA container toolkit and a host driver supporting the
chosen CUDA version (cu126 requires R560+; override with TORCH_CUDA=cu130).
# Build and run a local CUDA image (TORCH_CUDA defaults to cu126)
make build-cuda
docker run --rm --gpus all -p 8000:8000 -e API_KEYS=key1 -e DEVICE=cuda -v hf-cache:/data/hf laya-api:cuda
# Or use the Compose override (adds the GPU device reservation and DEVICE=cuda)
make up-cudaBake checkpoints into the CUDA image with make build-model-cuda MODEL=english.
For a raw build, pass the args directly:
docker build \
--build-arg TORCH_INDEX=https://download.pytorch.org/whl/cu126 \
--build-arg TORCH_DISABLE_NATIVE_JIT=1 \
-t laya-api:cuda .The workflow builds the CUDA image on pull requests as a check, and publishes the
linux/amd64 variants (cuda, english-cuda, multilingual-cuda,
typed-decisions-cuda, all-cuda) alongside the CPU images on branches, tags and
manual runs. A GPU is not required to build the image, only to run it.
The full OpenAPI 3.1 schema is saved directly in the repository as
openapi.json (and served at /openapi.json). You can
generate type-safe API clients for TypeScript, Go, Python, etc.:
npx openapi-typescript ./openapi.json -o laya-client.d.ts# Generate Python SDK
npx @openapitools/openapi-generator-cli generate -i openapi.json -g python -o ./clients/python
# Generate Go SDK
npx @openapitools/openapi-generator-cli generate -i openapi.json -g go -o ./clients/goTo regenerate the schema file at any time:
make openapiThe project requires Python 3.14+ and uses uv for dependencies:
uv sync --group dev # install runtime and development dependencies
make test # service and HTTP contract tests
make check # ruff, ruff format, mypy (strict) and pyrightThe test suite substitutes the Laya boundary, so it never downloads checkpoints or
runs inference. The CI workflow also runs a container smoke test (/healthz +
/predict) before publishing.
# Build and run the container
docker build -t laya-api:test .
docker run --rm -p 8000:8000 -e API_KEYS=test laya-api:testmake openapi # regenerate openapi.json
make test # pytest
make run # uvicorn --reload on :8000
make format / check / fix # ruff, format and static checks
make up / down / build / logs # docker compose
make build-cuda # CUDA image (adds the CUDA torch build args)
make up-cuda # docker compose with the CUDA overrideapp/
├── main.py # ASGI entry point: app.main:app
├── application.py # app factory and composition root
├── api.py # HTTP routes
├── authentication.py # API key, bearer and Basic auth
├── backend.py # typed boundary to the Laya library
├── config.py # per-instance immutable settings
├── schemas.py # Pydantic request/response contracts
└── services.py # runtime, prediction and SystemOne adapter
examples/ # HTTP requests and the category client
tests/ # service and HTTP contract tests
Laya is imported lazily, so openapi.json can be generated and validated without
torch installed.
- Runs as non-root (
appuser, uid 10001); thehf-cachevolume inherits that ownership. If you bind-mount your own cache, make sure uid 10001 can write it. - Torch uses the CPU wheel index (
pyproject.toml); ~0.35s per predict on CPU. Build withTORCH_INDEX=.../cu126to swap in the CUDA wheel for GPU inference. - Checkpoints load inside
transformers.initialization.no_init_weights()and the glibc heap is trimmed afterwards, which avoids the >3.7 GB peak of the multilingual encoder and returns freed load buffers to the OS. - The router loads once behind a lock, so requests are serialised per process.
/predict/bulkloops (laya's public API is single-state); it does not batch.
MIT © chneau