Skip to main content

Add captions

Captions are short descriptions used by selection and sharing checks. They are not the date and place labels printed on the film.

Basic works without them. GPU and Full describe the selected pictures and replacement candidates, then reuse those descriptions on later cuts. You do not need to caption the whole library before making a film.

The default captioner is SmolVLM2 500M. It receives small picture previews. Keep the service on your private network.

One NVIDIA worker​

If you are adding GPU inference and rendering too, use the one-GPU setup. It serves captions at /v1 in the same container, starting its bundled caption process on demand. Set caption_base_url to http://gpu-box:8092/v1. No separate caption container is needed.

The routes are unauthenticated. Keep the worker private. The GPU and Full tiers still need Laya in the app; captions alone do not enable GPU selection.

Separate Docker Compose captioner​

The released GPU file starts inference, the caption weight downloader and the caption server:

docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d
curl -s localhost:8094/v1/models

The preset supplies the inference and caption service URLs below saved Settings. On an existing install, change those URLs in Settings. The captioner is published on host port 8094; inside Compose it listens on 8092. GPU is a requested tier: preflight still checks the inference provider and Laya runtime. A CPU service does not satisfy GPU readiness.

For NVIDIA, use the CUDA image and device reservation together. The exact CUDA recipe and Kubernetes overlays are in the service reference.

Apple Silicon​

Use the mlxcel recipe to serve the pinned model natively. Point the app at it:

advanced:
editorial:
preparation:
caption_base_url: http://localhost:8092/v1

From Docker Desktop, replace localhost with host.docker.internal. A Docker container cannot use the Mac’s Metal GPU directly; configure GPU inference separately.

Check it​

immich-memories preflight

Look for Captions OK Serving smolvlm2-500m-base-public. The app checks the served model and synthetic control pictures before sending your library’s pictures. GPU and Full need a working caption provider; Basic with default settings skips it.

The service reference explains unreachable, wrong-model and authentication failures.

Explicit LLM captions​

A reader configured for prose never receives pictures automatically. To send previews and video frame strips to a vision-capable LLM instead, opt in:

advanced:
editorial:
preparation:
caption_provider: llm

This uses your configured LLM and its credentials. It can cost much more than the small local captioner, especially hosted. The app warns before using it. Existing valid SmolVLM results are reused first.

Model pins, serving flags, caption provenance and motion-description contracts are in the caption service reference.