Skip to main content

Run inference on another machine

Move picture analysis off a slow NAS. On NVIDIA, one CUDA worker can also caption pictures, split music into stems and render the film. Start with the app alone; add this when preparation or rendering takes too long.

One NVIDIA container​

Use the one-GPU setup. It serves these URLs from one container on port 8092:

App settingWorker URLWork
advanced.inference.facts_base_urlhttp://gpu-box:8092Picture classifiers and Demucs stems
advanced.editorial.preparation.caption_base_urlhttp://gpu-box:8092/v1SmolVLM captions
render.worker_base_urlhttp://gpu-box:8092/renderVideo rendering

The worker switches between classifier, caption, audio and render phases. It unloads classifier weights or stops its caption process when another phase needs the GPU. It keeps existing queues within each phase; it does not run all the models at once or promise that every card will fit them. The text reader, Laya family-viewing check and ACE-Step generation remain outside this container.

Only /render requires the render bearer token. Classifiers, captions and stems are unauthenticated. Keep the entire address on a trusted private network. The render worker also receives your Immich API key to download originals.

Classifiers and stems only​

The GPU tier file starts the inference and caption services. Its default CPU images are useful for diagnosis, but GPU readiness needs CUDA-capable inference:

docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d
curl -s http://localhost:8092/health

Connect the app using the shipped Compose service name:

advanced:
inference:
facts_base_url: http://immich-memories-inference:8092
fallback_to_local: true

From another machine, use its private LAN address. For NVIDIA standalone inference, select the CUDA image and its GPU device reservation; deployment recipes show both. This profile does not start the unified worker or a caption server.

Check it​

immich-memories preflight
curl -s http://gpu-box:8092/health

Health names each loaded producer's execution provider; empty lists mean it has not loaded yet. Look for CUDAExecutionProvider when you expect NVIDIA. A CUDA image can fall back to CPU, so an answering port alone does not prove acceleration.

With tier: auto, a GPU inference service, a working local CUDA ONNX runtime, or a Mac's Metal GPU enables GPU; an enabled reader with a model makes that Full. Captions and Laya still need to be ready. Hardware video encoding alone does not change selection tier.

Keep caches on volumes. Completed matching facts stay in the app's store when you move the service. fallback_to_local: true lets the app attempt missing facts or stems locally if its runtime and model files are installed. Set it to false when a failed service should stop the cut.

Service reference: artifact paths, queue limits, offline setup and memory ownership. Render setup: tokens, transport and output checks.