Skip to main content

Inference service deployment and API

On Basic the app runs the DINOv2 encoder and its eight heads in its own process, on the CPU, once per picture. GPU and Full also enable Marqo and Docling. The inference service moves the active producers to another machine: a GPU box, a Kubernetes node, or just a container you can restart on its own. The facts are the same rows either way, so you can add it, move it or drop it without re-deriving anything.

For NVIDIA, use the one-GPU setup to serve classifiers, SmolVLM captions, Demucs and rendering in one container. The ordinary inference entry point serves classifiers and stems only. Its Compose and Kubernetes recipes below remain separate from caption and render services.

Both modes use port 8092. In unified mode, captions are under /v1 and authenticated rendering under /render. In standalone mode, a separate caption server needs its own host port; shipped Compose publishes it on 8094.

Picture previews and full music tracks reach inference. Unified rendering also receives your Immich API key. Only /render authenticates requests: keep the listener private. ACE-Step and the text reader are not served here; Laya remains in the app process.

Unified worker admission and memory​

The CUDA image's default command is still python -m immich_memories_inference. To start all four services, use python -m immich_memories_inference.gpu_worker through the shipped services/inference/compose.gpu-worker.yaml recipe. It requires the shared render token, the app's Immich URL and a scratch directory; the setup guide shows the environment and app settings.

Requests acquire a classifier, caption or audio phase. A different active phase returns HTTP 503 with Retry-After: 1. Render work waits up to 60 seconds for model work, then owns the GPU; model requests are refused while it is preparing or rendering. Cancellation does not release active native work before that work finishes. This is phase admission, not one global FIFO.

Before a phase change, the worker unloads classifiers when leaving facts and stops its caption child when leaving captions. Demucs releases after its response. The next caption request restarts the child, which listens only on loopback inside the container. Render authentication headers are not forwarded to it.

The bundled caption command sets context to 8192 tokens but does not explicitly set a RAM prompt-cache limit or processing-slot count. Phase changes reduce overlapping model residency; they do not impose a hard host-RAM or VRAM budget. The Compose recipe has no RAM limit. Other containers and external readers/music backends can still consume the same card.

Verify the rendering runtime​

For the unified worker, keep the /render suffix in render.worker_base_url. Check http://WORKER:8092/render/health for render readiness; /health checks inference and does not establish that the render route has the expected contract or acceleration support.

Use the app and worker images from the same release. When developing with source mounted over an older image, install that source revision's declared dependencies too. A source copy does not update its Python environment. Older custom images may lack pi-heif: inference succeeds, then remote assembly rejects HEIC photos.

Before a long run with a custom image, check import pi_heif in the worker's Python environment and decode a representative HEIC source. A successful health response does not exercise every media decoder.

The two images​

TagPlatformsProvider
:X.Y.Zlinux/amd64, linux/arm64CPU
:X.Y.Z-cudalinux/amd64CUDA, falling back to CPU

openvino, armnn and rocm are not shipped. Quick Sync, VAAPI and NVENC decode, scale and encode; they run no inference.

The CUDA image uses ONNX Runtime GPU for the DINOv2 encoder, Marqo and Docling. The eight small heads project the encoder output with NumPy. A graph rejected by CUDA falls back to CPU; the image tag alone does not certify GPU execution.

Use the published release image tag matching your app. The image names are:

docker pull ghcr.io/sam-dumont/immich-memories/inference:YOUR_APP_TAG
docker pull ghcr.io/sam-dumont/immich-memories/inference:YOUR_APP_TAG-cuda

From a checkout, docker/Dockerfile.inference builds either one: --build-arg DEVICE=cpu or DEVICE=cuda, with --build-arg APP_VERSION=0+local.

Music stems​

With advanced.inference.facts_base_url set, the app sends generated music to POST /audio/stems and receives a ZIP containing drums.wav, bass.wav, other.wav and vocals.wav. The endpoint accepts one multipart file, up to 256 MiB. One separation runs at a time, off the HTTP event loop: an upload that arrives while another separates gets HTTP 429 with Retry-After before its body is read, and the app falls back as for any failed stem request. Temporary audio is deleted after the response. Both inference images include Demucs. The CUDA image bundles its weights; the CPU image downloads them on first use into /cache/torch.

advanced.inference.fallback_to_local also controls recovery from a failed stem request. Its default is true: the app uses local Demucs if installed. An explicitly enabled MusicGen server keeps priority for stems. Neither setting changes the ACE-Step generation endpoint.

Running it with compose​

The base file runs Basic. Add the released GPU tier file to start inference and captions:

docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d
curl -s localhost:8092/health

For NVIDIA, install the NVIDIA container toolkit and use Linux driver 570.124.06 or newer as the recommended baseline for the image's CUDA 12.8.1 runtime (NVIDIA driver table). CUDA 12 minor compatibility can run on drivers from 525.60.13, with feature and PTX/JIT limits; that lower floor does not guarantee this image's kernels work. Newer cards may need a newer driver. Check that a container sees the card:

sudo nvidia-ctk runtime configure --runtime=docker && sudo systemctl restart docker
docker run --rm --gpus all nvidia/cuda:12.8.1-base-ubuntu24.04 nvidia-smi

Add the CUDA file; it selects the CUDA images and reserves the device together:

docker compose -f docker-compose.yml -f docker-compose.gpu.yml -f docker-compose.cuda.yml up -d
curl -s localhost:8092/health

Set IMMICH_MEMORIES_VERSION once in .env for the app and inference image. There is no separate inference tag. The GPU preset requests GPU; a CPU image does not satisfy its inference compute check. Change the inference URL in Settings for an existing install, or use the setup builder for fresh-install defaults that remain editable there.

The app detects remote GPU capability from the top-level /health fields status: "ok" and provider: "CUDAExecutionProvider"; that probe loads no models and sends no pictures. /health also names the provider per producer, in producers.heads.providers, producers.nsfw_marqo.providers and producers.doc_docling.providers, each empty until that producer has loaded. CPUExecutionProvider on a GPU host means the image is the CPU one, the device reservation did not reach the container, PROVIDER names cpu, or the driver and the CUDA runtime in the image do not match. Check the tag first: it is the usual one. A graph ONNX Runtime turns down gets one WARNING naming the seat and the fallback, so the service never serves CPU answers quietly.

On Kubernetes​

deploy/kubernetes/overlays/inference is the service on its own: a Deployment, a ClusterIP Service on 8092, a 10Gi model-cache PVC and a NetworkPolicy. It does not pull in base/, so it needs no Secret and no Immich, and runs in a cluster where the app does not.

kubectl create namespace immich-memories # base/ creates it too
kubectl apply -k deploy/kubernetes/overlays/inference # CPU
kubectl apply -k deploy/kubernetes/overlays/inference-cuda # NVIDIA nodes

inference-cuda is the same overlay plus one patch: runtimeClassName: nvidia, one nvidia.com/gpu, the two NVIDIA_* env vars, the nvidia.com/gpu.present=true node selector, the matching toleration and the -cuda tag. Each overlay pins its own tag in an images: entry; bump both together and match them to the app version you are deploying.

Port-forward and read the provider back:

kubectl -n immich-memories port-forward svc/inference 8092:8092
curl -s localhost:8092/health

Then point the app at it: http://inference:8092 in the same namespace, http://inference.immich-memories.svc.cluster.local:8092 across namespaces. The base NetworkPolicy already allows the app egress on 8092.

A cold cache volume​

The CUDA image bundles DINOv2, the context heads, Marqo, Docling, Laya ONNX and the SmolVLM2 caption model and projector. Its weights live in /opt/immich-models, outside the writable cache mount. Downloads are unnecessary at startup; the image sets HF_HUB_OFFLINE=1. The app's CLI can use the same image, with the bundled paths already configured.

For standalone inference, a separate caption container can reuse the CUDA image and layers. The unified worker starts its own child; do not add this container to that setup:

image=ghcr.io/sam-dumont/immich-memories/inference:YOUR_APP_TAG-cuda
docker run --rm --gpus all -p 127.0.0.1:8094:8092 \
-v immich-memories-model-cache:/cache \
"$image" immich-memories-captioner --cache-ram 128 --parallel 1

Point advanced.editorial.preparation.caption_base_url at http://localhost:8094/v1 when running the app on the host. Between containers, use the caption container's hostname and port 8092. The llama.cpp runtime bundled in the image is pinned by digest; older NVIDIA cards may compile kernels on their first request. Its bounded JIT cache stays on /cache across restarts. Laya runs in the app process; the /facts service serves the image classifiers.

The CPU image keeps its smaller download. A fresh PVC is empty, which is fine. Both overlays set ALLOW_MODEL_DOWNLOADS=true, and the CPU service then fetches what it is missing on first use: the pinned DINOv2 export (88 MB), the pinned Marqo export (22.5 MB) and the Docling snapshot. The ONNX exports are checked against the same SHA-256 immich-memories models fetch pins, the Docling snapshot by Hugging Face revision. Only the first /facts call after a cold start waits for it.

For an offline CPU deployment, provision /cache/dinov2-small.onnx, /cache/nsfw-marqo-384.onnx and the pinned Docling snapshot under /cache/huggingface before starting the service. Copy the complete Hugging Face cache, including snapshot metadata. The model download command needs --detectors to fetch Marqo and Docling on Basic; its default destination is not the service's /cache volume. Match the service's configured paths when copying artifacts. With ALLOW_MODEL_DOWNLOADS=false, a missing model returns 503 naming the artifact and expected path.

Offline CPU stem separation additionally needs the htdemucs Torch checkpoint cache under /cache/torch. ALLOW_MODEL_DOWNLOADS controls classifier downloads, not Demucs. Provision that cache before disconnecting the service, or use the CUDA image with bundled stem weights.

The pod's root filesystem is read-only, so the overlay sets HF_HOME=/cache/huggingface and TMPDIR=/tmp. Without them the download has nowhere to put its temporary files, fails with Read-only file system (os error 30), and every /facts request answers 503 until the pod restarts. Keep both if you write your own manifests.

Reach the service from outside the cluster​

http://inference:8092 only resolves inside the cluster. For a NAS or a laptop on the LAN, overlays/inference-lan adds a second Service, type LoadBalancer, on the same pods and port. The ClusterIP Service is untouched, so no in-cluster caller starts riding an external address, and it composes with either device overlay.

kubectl apply -k deploy/kubernetes/overlays/inference-lan
kubectl -n immich-memories get service inference-lan

Put what that prints in facts_base_url. The address comes from the cluster's load-balancer controller: without one the Service sits at <pending> forever. Nothing behind port 8092 checks a credential, so do not give it a routable address, and kubectl delete -k deploy/kubernetes/overlays/inference-lan when you are done.

What it answers​

EndpointQuestion
GET /pingare you up
GET /healthwhich producers are loaded, at which versions, on which provider each
GET /queuewaiting and active classifier work, completions, failures and timings
POST /factsone picture in: what do the frozen classifiers say about it
python3 -c 'import base64,json,sys; print(json.dumps({"image": base64.b64encode(open(sys.argv[1],"rb").read()).decode(), "producers": ["heads"]}))' photo.jpg \
| curl -s localhost:8092/facts -H 'content-type: application/json' --data-binary @-
{"producers": {"heads": {"encoder_key": "…", "facts": [
{"head": "activity", "version": "public-v1", "label": "outdoors", "confidence": 0.71}]}}}

When a producer cannot load, /facts answers 503 with the reason in detail: which producer, which artifact, which path. The app repeats that as Remote classifiers returned HTTP 503: <detail>, and the service logs it once per distinct message rather than once per picture. A failed load is retried on a later request, backing off from 10 s to 2 minutes, so fixing the cause needs no restart.

A fact on the wire is the bank row without its asset id, and the client stores it verbatim: a fact's identity is the artifact that produced it, never the machine that ran it. No URL, hostname, device or provider name enters any key, so the same picture through the cpu and the cuda image lands on one row, label-identical rather than byte-identical.

Watching classifier work​

curl -s localhost:8092/queue

The response reports queued, active, completed, failed, cancelled and rejected for each producer, plus oldest_wait_seconds, mean_wait_seconds, mean_run_seconds and completed_per_second. Counts and averages cover this service process since startup; throughput includes idle time. Waiting time includes admission to the shared worker pool. Model loading counts as run time. The response contains no pictures or asset identifiers.

Each model has a FIFO queue. Waiting callers yield the worker so another model can run. At most REQUEST_THREADS model calls run at once, with one active call per model. The service accepts 32 waiting calls across the classifier queues by default. A full queue returns HTTP 429 with Retry-After: 1; callers should reduce concurrency and retry later. The app's configured local fallback still applies if the request fails. Cancelling a waiting call removes it; cancelling active native work keeps its worker reserved until that work finishes.

These queues cover /facts, in both entry points. Standalone Demucs has its own serial scheduling. Unified mode adds phase admission across facts, captions, audio and rendering; /queue does not report those other phases. Separate caption servers based on llama.cpp can expose slots and metrics when their server settings enable them.

Request limits​

The service checks each request's size before reading its body:

RouteLargest bodyOver the limit
POST /factsthe base64 of MAX_IMAGE_BYTES (about 21 MiB by default), plus 1 MiB of JSONHTTP 413
POST /audio/stems256 MiB of audio, plus 1 MiB of multipart envelopeHTTP 413

A declared Content-Length over the limit is refused unread. A body that runs past it, declared or not, is cut off with the same 413. /facts holds at most REQUEST_THREADS plus MAX_QUEUED_REQUESTS bodies at once; one more gets HTTP 429 with Retry-After.

A picture whose header asks for more than 50 million pixels gets HTTP 413 naming its pixel count, before it is decoded. The file size says little here: a few kilobytes of PNG can decode to gigabytes.

Settings​

Every setting is an environment variable prefixed IMMICH_MEMORIES_INFERENCE_:

VariableDefaultWhat it does
HOST / PORT127.0.0.1 / 8092where to listen. The image sets the host to 0.0.0.0
CACHE_DIR/cachethe model cache volume
ENCODER$CACHE_DIR/dinov2-small.onnxthe pinned DINOv2 export, digest-verified on load
MARQO_ONNX$CACHE_DIR/nsfw-marqo-384.onnxthe pinned sensitive-content ONNX export
BUNDLEthe packaged public bundlehead bundle .npz
PROVIDERautoauto, cpu, cuda or coreml. auto takes CUDA where the provider is present and CPU otherwise. CoreML is selectable; Linux images use CPU or CUDA
REQUEST_THREADS4the thread pool in front of ONNX Runtime. The app's facts_concurrency is what fills it
MAX_QUEUED_REQUESTS32maximum waiting calls across classifier queues; excess requests get HTTP 429
IDLE_UNLOAD_SECONDS300drop idle weights; 0 holds them
PRELOADfalseload every producer at boot instead of on first use
DETECTOR_CACHE_DIRunset in the runtime; /opt/immich-models/huggingface in the CUDA imagewhere the detector snapshots live; unset uses the Hugging Face cache, with HF_HOME=/cache/huggingface in the CPU image
ALLOW_MODEL_DOWNLOADSfalselet a cold cache fetch the pinned exports and the Docling snapshot itself
MAX_IMAGE_BYTES16777216refuse a decoded picture larger than this, and size the /facts body limit from it

The CUDA image sets OPENBLAS_NUM_THREADS=1 for NumPy's small per-picture head projection. This limits NumPy's BLAS worker pool, which otherwise can compete for a container's CPU quota. REQUEST_THREADS independently controls classifier request concurrency. Measure under your actual CPU and memory limits before overriding either setting.

Idle unload drops the weights and keeps the process: the next request reloads them. Changing the provider re-keys nothing, so any of this can be retried without re-deriving a fact.

Point the app at it​

Use the app connection and health checks. The configuration reference lists timeout, concurrency and fallback settings.