Skip to main content

Caption service contract

The caption setup guide gets the Compose service running. This reference keeps explicit LLM captioning, pinned artifacts, alternative deployments and the contract every server must meet.

Unified or separate service​

The unified NVIDIA worker already exposes captions at /v1. Its phase manager stops the owned caption subprocess between other GPU phases and restarts it on demand. The standalone Compose/Kubernetes recipes below run their own server; do not add one just to use unified captions.

The unified caption route has no authentication, even though /render on the same listener requires a token. Keep it private. The standalone recipes explicitly set --cache-ram 128 and --parallel 1; the unified bundled command does not set those bounds. Context size is 8192 in both llama.cpp recipes. Neither configuration guarantees a fixed VRAM peak.

Explicit LLM captions​

SmolVLM is the default caption provider. An LLM configured for titles or prose never receives images automatically. To let a vision-capable LLM supply missing captions, opt in:

advanced:
llm:
provider: openai-compatible
base_url: http://localhost:8000/v1
model: your-vision-capable-model
editorial:
preparation:
caption_provider: llm

This is less efficient than SmolVLM and can be much more expensive, especially on hosted infrastructure. The app warns at startup and in preflight. Image tiles and video frame strips go to the configured LLM, using its credentials and provider settings. The SmolVLM endpoint and caption_api_key are unused for this choice.

Existing valid SmolVLM captions and motion lines stay banked and are reused first. New LLM captions have a separate producer identity and provenance; changing models does not relabel old captions. Before new picture captions, three synthetic tiles must pass the schema check. A failed LLM caption stays outstanding, with no automatic fallback to another model.

Film generation still captions only the selected shots and actual candidates. prepare is the explicit job for a wider scope. This choice works on Basic without promoting selection to Full: automatic tiers still follow available GPU inference capability. Sharing decisions remain with the rules, classifiers and Laya where enabled.

The SmolVLM contract​

Before a single library picture goes on the wire, the app checks that GET /models advertises smolvlm2-500m-base-public, then sends three synthetic control tiles (red, blue, grey) and requires a schema-valid answer to each. Then each picture is one 400 px JPEG tile, one request at temperature 0 with a repetition penalty of 1.1, a 140-token cap and a JSON schema the server has to honour. Two invalid answers bank caption unavailable; a timeout, a missing model or a transport error stays outstanding and the next run picks it up.

The alias is a promise about behaviour, not a name lookup: any endpoint can claim it. It means the descriptions under that name came from SmolVLM2-500M with this prompt and this schema, so a bank filled last month and one filled today are comparable. Alias an unrelated vision model and the app will believe you, and the bank then holds two things under one name. The default SmolVLM provider expects this alias from your server. The explicit LLM option above uses the configured model's own identity instead.

caption_api_key goes out as Authorization: Bearer <key>; blank sends no header. The reader's llm.api_key is never borrowed for it.

Caption reports name the producer of the caption actually reused, which can differ from the currently configured model. A description and its setting are read as one complete pair from the same producer.

Accepted artifacts​

FormatRepositoryRevisionRuns on
MLXmlx-community/SmolVLM2-500M-Video-Instruct-mlxfa57db46815177fbdfd65cc85a2b3416a8332268Apple Silicon
GGUFggml-org/SmolVLM2-500M-Video-Instruct-GGUFccd7aae53bcb1997355c2f094959e72b3642ce17Anything llama.cpp runs on

Same 500M model either way; pick what your hardware runs. The GGUF files this page pins:

FileSizeSHA-256
SmolVLM2-500M-Video-Instruct-Q8_0.gguf437 MB6f67b8036b2469fcd71728702720c6b51aebd759b78137a8120733b4d66438bc
mmproj-SmolVLM2-500M-Video-Instruct-Q8_0.gguf109 MB921dc7e259f308e5b027111fa185efcbf33db13f6e35749ddf7f5cdb60ef520b

Use the pinned Q8_0 artifacts to match the bundled serving contract.

Apple Silicon, with mlxcel​

A native Apple Silicon serving option:

brew install lablup/tap/mlxcel
pip install huggingface-hub # for the `hf` command, if you do not have it
SNAPSHOT=$(hf download mlx-community/SmolVLM2-500M-Video-Instruct-mlx \
--revision fa57db46815177fbdfd65cc85a2b3416a8332268)
mlxcel serve --model "$SNAPSHOT" --alias smolvlm2-500m-base-public --host 0.0.0.0 --port 8092

--host 0.0.0.0 because mlxcel binds 127.0.0.1 by default, which an app on a NAS or another host cannot reach (Caption endpoint unreachable); nothing behind the port checks a credential, so keep it on your LAN.

hf download prints the snapshot directory it wrote, which is what the server wants. Then in your config:

tier: auto
advanced:
editorial:
preparation:
caption_base_url: http://localhost:8092/v1

That is the default value of caption_base_url. On an Apple Silicon Mac installed with the all-mac extra, automatic selection sees the Metal GPU and chooses GPU, or Full when an LLM is also configured. Install the Laya checkpoint too. From the app in Docker Desktop on the same Mac, the address is http://host.docker.internal:8092/v1; from a NAS, the Mac's LAN name or IP. If you run oMLX for the reader, use the separate SmolVLM caption server with its own process on its own port. A container does not see the Mac's GPU: configure a GPU inference service as well. A caption URL alone does not select the GPU tier.

Apple Silicon, with llama.cpp​

Switching the caption server from mlxcel to llama.cpp cuts its resident memory by about 4x on a 16 GiB Mac; see Measure your setup for the exact figures. These are measured service footprints, not whole-app memory requirements. The reader and music still need their own memory.

Use the llama-server executable installed for the local reader and the GGUF files pinned in the table above. This is an alternative to the mlxcel caption process, so stop that process before using its port; leaving both loaded defeats the memory saving. Keep the same served alias.

CAPTION_MODELS="$HOME/.immich-memories/models/captioner-gguf"
hf download ggml-org/SmolVLM2-500M-Video-Instruct-GGUF \
SmolVLM2-500M-Video-Instruct-Q8_0.gguf \
mmproj-SmolVLM2-500M-Video-Instruct-Q8_0.gguf \
--revision ccd7aae53bcb1997355c2f094959e72b3642ce17 \
--local-dir "$CAPTION_MODELS"
llama-server \
--model "$CAPTION_MODELS/SmolVLM2-500M-Video-Instruct-Q8_0.gguf" \
--mmproj "$CAPTION_MODELS/mmproj-SmolVLM2-500M-Video-Instruct-Q8_0.gguf" \
--alias smolvlm2-500m-base-public --host 127.0.0.1 --port 8092 \
--ctx-size 8192 --cache-ram 128 --parallel 1 --threads 4 --n-gpu-layers 99

Reuse these files if they are already installed, checking their hashes against the pinned table. This recipe was checked with llama.cpp build 11146; if captions fail to load, try the latest release. Point advanced.editorial.preparation.caption_base_url at http://127.0.0.1:8092/v1 and run preflight. The application checks the real caption schema before sending missing library previews. A passing caption check does not certify the whole selection/render/music handoff.

Docker and Linux, with llama.cpp​

The released GPU tier file includes a weight downloader and a caption server. The downloader checks both pinned digests before starting the server:

docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d
curl -s localhost:8094/v1/models

The first start downloads about 546 MB of model and projector weights, plus container image layers. Subsequent starts verify the cached weights. Compose publishes captions on host port 8094 and inference on 8092. Inside the network, both use port 8092. The GPU preset supplies those service URLs as defaults below saved Settings. Existing installs can change the inference and caption URLs in Settings; runtime URL environment overrides would pin them above Settings.

GPU is a requested tier, not proof of a working device. Run preflight after startup. CPU inference does not pass the GPU compute check. Preparation follows the product tier; do not set a separate preparation tier.

Running ghcr.io/ggml-org/llama.cpp:server by hand works the same way, with the weights bind-mounted at /models and --host 0.0.0.0 --ctx-size 8192 --cache-ram 128 --parallel 1 --threads 4. The RAM prompt cache is bounded at 128 MiB, with one processing slot. These are host-RAM prompt-cache and processing-slot limits, not a total RAM or VRAM ceiling. Measure them alongside other GPU users. Completed captions stay in the app's persistent store.

Three more flags carry the serving contract:

FlagLeave it out and
--alias smolvlm2-500m-base-publicthe server advertises the GGUF path instead, and preflight says the endpoint serves another model
--mmproj …the model loads and answers, but cannot interpret images: the control tiles fail and no library picture is sent
--port 8092nothing answers where the app looks, and the row reads unreachable

Use local model and projector files with the pinned runtime. Preflight checks /models and the advertised alias. Preparation sends synthetic control tiles first; failed controls stop it before any library picture is sent. Preflight does not validate image responses.

On an NVIDIA host​

Add the released CUDA file to the base and GPU files:

docker compose -f docker-compose.yml -f docker-compose.gpu.yml -f docker-compose.cuda.yml up -d

It selects the pinned server-cuda-b10920 caption runtime, requests the NVIDIA device and sets LLAMA_ARG_N_GPU_LAYERS=99. No separate captioner version variable or hand-written override is needed. The value 99 offloads all 32 layers of the pinned 500M model. Install the NVIDIA driver and container toolkit first, then check captions and inference with preflight.

Optional app concurrency, after measuring the server under load:

IMMICH_MEMORIES_EDITORIAL__PREPARATION__CAPTION_CONCURRENCY: "4"

Speed, and what to set caption_concurrency to​

Start with the default of 1. More concurrent requests can make a CPU server slower because they compete for the same threads. On a GPU, try 4 and measure with the other services you run. Report new captions separately from cache hits: a selection run that reuses captions does not measure caption throughput. Dated measurements belong on Measured.

Kubernetes​

kubectl create namespace immich-memories # if you have not already
kubectl apply -k deploy/kubernetes/overlays/captioner # CPU
kubectl apply -k deploy/kubernetes/overlays/captioner-cuda # NVIDIA nodes

A Deployment, a ClusterIP Service on 8092, a NetworkPolicy and a 2 Gi PVC an init container fills and digest-checks before the server starts. Neither overlay includes base, so both apply without the Immich secret: the captioner holds no credential. It does receive a 400 px tile of every picture it is asked to caption, so the Service stays ClusterIP and the NetworkPolicy allows ingress on 8092 only.

Point the app at it:

tier: auto
advanced:
inference:
facts_base_url: http://inference:8092
editorial:
preparation:
caption_base_url: http://captioner:8092/v1

Across namespaces that is captioner.immich-memories.svc.cluster.local:8092.

The app's base manifest uses tier: auto. Point it at the GPU inference and caption services; the resolved product tier sets preparation too:

kubectl -n immich-memories set env deployment/immich-memories \
IMMICH_MEMORIES_INFERENCE__FACTS_BASE_URL=http://inference:8092 \
IMMICH_MEMORIES_EDITORIAL__PREPARATION__CAPTION_BASE_URL=http://captioner:8092/v1

captioner-cuda is the same Deployment with the server-cuda-b10920 image, --n-gpu-layers 99 appended, and the three things the GPU Operator wants: runtimeClassName: nvidia, the nvidia.com/gpu.present node selector and the matching toleration. Nothing else changes, so pointing the app at it is the block above plus caption_concurrency: 4.

It deliberately does not request nvidia.com/gpu: 1. Where one card is time-sliced per node that resource has a single slot, the inference Deployment holds it, and a captioner asking for a second stays Pending beside an idle card. Without the request, it relies on the NVIDIA runtime exposing the same card. This is not a portable GPU-sharing or memory-budget guarantee. The unified worker avoids separate inference and caption Deployments; use its Compose recipe when one service fits your deployment. With a card to spare, put the request back:

resources:
limits:
nvidia.com/gpu: "1"
requests:
nvidia.com/gpu: "1"

The standalone overlays pin llama.cpp build b10920: server-b10920 on CPU and server-cuda-b10920 on CUDA. Update both together and validate the serving contract.

How preflight reports it​

Captions OK Serving smolvlm2-500m-base-public is the row you want. Failures name the cause:

RowWhat happened
Caption endpoint unreachablenothing is listening, or it is not HTTP
Caption server is slow to answerthe model inventory exceeded editorial.preparation.caption_timeout_seconds (90 seconds by default); check the worker logs or raise that timeout for cold model loading
Caption endpoint serves another modela server answered and advertised something else
Caption endpoint refused the request401 or 403, so set caption_api_key

On Basic with the default caption provider the row reads SKIPPED. An explicit LLM-caption opt-in prints the image-sharing warning; it does not send a vision request. Preparation validates that provider's synthetic control responses before sending library pictures.

What a missing captioner costs​

On GPU and Full with SmolVLM, prepare stops: the description producer stays outstanding and the failure names caption_base_url. Default Basic does not require that endpoint. With explicit LLM captions, a failed image request stays outstanding under that provider's identity. What each tier runs: Requirements and tiers.

Knowing which build wrote a caption​

The alias is a contract, not a build identifier: both recipes on this page advertise it on port 8092. So each new caption row also keeps the server's /models row and a 16-character digest of its answers to the three controls, plus an optional label of your own (advanced.editorial.preparation.caption_artifact_id). prepare prints one line naming every captioner behind the bank, with MIXED when there is more than one, and immich-memories runs why <asset-id> --run <run-id> shows the one a run used. None of it re-captions anything already banked.

Switching servers later​

The bank keys on the producer name, description:smolvlm2-500m-base-public@envelope-v3-compact, which carries no format, no quantisation and no weights digest. Swapping MLX for GGUF re-captions nothing: every banked picture keeps the wording the old server gave it, and short of clearing the description rows there is no way to ask for a re-caption.

Different runtimes can phrase the same picture differently. A mixed bank keeps each result's provenance, but switching servers does not normalise existing wording. Pick one if consistency matters for your library.

Motion lines​

On the gpu and full tiers, every video a cut prepares gets one banked sentence about what happens in it, from the same caption server and model. So does every Live Photo whose motion an earlier cut measured at 1.5 or more; one nobody measured yet is not known to play and gets none. The story pick reads that sentence beside the video's row, so the reader compares a video with a still in text and never sees a video frame.

The app does not download the video for it. Immich answers byte ranges on its playback rendition, so preparation reads the index (tens of kilobytes), picks the three keyframes nearest a quarter, half and three quarters of the clip, and reads only those. FFmpeg copies exactly those packets out of a sparse local copy and decodes them, which behaves the same on FFmpeg 5.1 (the app image), 6.1, 7.1 and 8.1. A codec other than H.264, HEVC, VP9 or AV1 also costs its first keyframe, which FFmpeg needs to read the stream at all. The three frames go to the server as one 960 × 320 JPEG strip with a one-field schema (description, 120 characters), under the same temperature, penalty, token cap and caption_api_key as captions.

The bank keys on the picture, its complete source metadata and motion-line-v1@smolvlm2-500m-base-public/3-keyframes-320px, so a changed source is asked again and nothing else is. Two invalid answers, a playback Immich answers 404 for, or an index the app cannot read are banked as settled. Timeouts and transport failures stay missing, so a later prepare retries them and reports a visible "motion unavailable" warning. Normal film refinement reuses its rules draft and does not consume these sentences, so it does not request them or report their absence as a preparation failure. A model-planning route can use banked lines or plain clip facts. caption_concurrency bounds the requests in flight; keyframe reads run four at a time.

Each row also records what produced it: a digest of the question asked, the keyframe times it read, and what made the source owe a line (video, or the Live Photo's residual and the measurement that produced it). Missing provenance is reported as unrecorded in the motion metrics.

A video's sentence counts only where its motion is measured. On every tier that samples a video's frames for the exposure head, preparation also measures their optical-flow residual and banks it under the picture, its source metadata and motion-residual-v1@median-flow-v1-detector-frames-320x240, with the frame count it was measured on. A video that measures under 1.5 has its sentence withheld (a favourite keeps it); see Picking each shot. A missing residual can require another sample; a frame read or measurement failure is named and does not block the cut.

prepare owns motion-description requests. It banks true videos and Live Photo companions whose measured residual already qualifies them. A first cut can discover motion in an unmeasured Live Photo; it uses plain facts until a later prepare banks that companion's sentence. Cuts read banked descriptions and never contact the motion-description server. Missing lines remain listed in the private preparation record. Required captions and safety facts still gate the cut.

The basic tier asks for no motion line. The pick then reads the video's plain facts instead: its length, and the measured motion of a Live Photo that has one.