Skip to main content

The reference setup

Two machines, every optional piece turned on: a Kubernetes cluster that runs the film-making pipeline around the clock, and a Mac that runs the same app locally with nothing else around it. Written up from a real deployment. Most installs want a fraction of this: read What a GPU or a model adds first, and come back here for the pieces worth adding.

Placeholders throughout: photos.example.com, 192.168.1.50, gpu-node-a, idp.example.com. Nothing on this page is a real hostname, IP or node name.

The two profiles​

Always-on server (Kubernetes)                    Laptop / workstation (the Mac)
────────────────────────────── ──────────────────────────────
gpu-node-a (the newer card, time-sliced) one machine, everything local
├── immich-memories pod
│ ├── app container immich-memories ui
│ └── render-worker (sidecar, loopback :8093) ├── oMLX: gemma-4-e4b-it-6bit
└── inference service (encoder, heads, │ (OpenAI-compatible, :9999)
detectors; ONNX Runtime CUDA, :8092) ├── mlxcel: SmolVLM caption
│ server (:8092, localhost)
gpu-node-b (a second, older card) └── ACE-Step 1.5, lib mode
└── captioner (llama.cpp server-cuda, :8092) (XL turbo, bf16, 4B LM)

Traefik (TLS) ── OIDC (Auth0) ── the app OIDC to the same IdP
the reader (LLM): an OpenAI-compatible server on the LAN, e.g. oMLX on the Mac

The server profile is the app's Deployment plus two GPU services, reached through deploy/kubernetes/overlays/maximalist or deploy/terraform/examples/maximalist. The Mac profile is a source checkout with ACE-Step installed beside it, plus two local servers (the reader and the caption server): see Generated music and Add a reader.

Feature → where it runs → config keys → hardware​

FeatureWhere it runsConfig keysHardware
App + UIthe Kubernetes pod, or the Mac directlytier: fullany CPU
Render workersidecar in the app pod, loopbackrender.worker_base_url, render.worker_tokengpu-node-a: NVENC h264/hevc
Picture reading (inference service)its own pod, on the newer cardadvanced.inference.facts_base_urlgpu-node-a: ONNX Runtime on CUDA; without it, the app pod's CPU
Caption servera second GPU node, cluster-onlyadvanced.editorial.preparation.caption_base_urlgpu-node-b: a Pascal card works (sm_61)
Caption server (Mac)mlxcel, localhostsame key, http://localhost:8092/v1Apple Silicon, Metal
Reader (LLM)the Mac, on the LANadvanced.llm.base_url, advanced.llm.provider, advanced.llm.modelApple Silicon running oMLX
MusicACE-Step 1.5 API (cluster) or lib mode (Mac)advanced.ace_step.mode, advanced.ace_step.api_urlcluster: any card; Mac: Apple Silicon, XL wants more RAM
Music stemsDemucs in the inference serviceadvanced.inference.facts_base_urlinference: CUDA or CPU; app/render fallback: CPU; Mac fallback: Metal
OIDC behind a proxythe reverse proxy + the appadvanced.auth.public_url, advanced.auth.trusted_proxies, advanced.server.secure_cookiesnone
Geocoding + map tilesnominatim.openstreetmap.org, server.arcgisonline.comnetwork.geocoding, network.map_tilesnone
Cache capsthe cache PVC / local diskcache.video_cache_max_size_gb, cache.thumbnail_cache_max_size_mbsized storage
Automationin-process timer, or a CronJob to /api/triggeradvanced.automation.enabled, advanced.automation.daily_atnone

Missing a piece from this table: drop the row and set tier to match, because a named tier never steps down on its own. tier: full refuses to load without advanced.llm.base_url and advanced.llm.model, and still asks for captions when no caption server answers; tier: auto picks gpu or full only once it finds GPU picture reading (a local CUDA runtime or the inference service), so without one it stays on the plain NAS tier. No ACE-Step means a bundled track. Without the inference service the app reads pictures on its own CPU: the same answers, a slower first cut.

The always-on server (Kubernetes)​

cd deploy/kubernetes
cp base/secret.yaml.example base/secret.yaml
cp overlays/render-sidecar/render-worker-secret.yaml.example overlays/render-sidecar/render-worker-secret.yaml
cp overlays/maximalist/maximalist-secret.yaml.example overlays/maximalist/maximalist-secret.yaml
vim base/secret.yaml overlays/render-sidecar/render-worker-secret.yaml \
overlays/maximalist/maximalist-secret.yaml overlays/maximalist/config-map.yaml
kubectl apply -k overlays/maximalist

overlays/maximalist/kustomization.yaml composes overlays/render-sidecar (which itself pulls in base), overlays/captioner-cuda and overlays/inference-cuda, rather than duplicating any of them. It keeps the base's 20Gi cache PVC, so it applies over an existing install. On top: a second init container that installs config.yaml, IMMICH_MEMORIES_TIER=full on the app container (the base sets auto, and an env var beats config.yaml), and a NetworkPolicy that lets the app reach the LLM and ACE-Step ports. Read Kubernetes first for the base layout this builds on.

Check the tier it really runs​

The tier in the file and the tier the pod runs can differ: the environment wins over config.yaml, and the base Deployment sets IMMICH_MEMORIES_TIER=auto. Ask the app, before and after any change:

kubectl -n immich-memories exec deploy/immich-memories -c immich-memories -- \
immich-memories config show | grep '│ tier'

The row names the tier and where it came from. full from env is this setup. nas from env means the pod has been cutting on the plain NAS tier, the caption server and the reader never asked; the log line above the table says why (Automatic selection tier: nas. ...).

The cluster's config.yaml, annotated​

The full file overlays/maximalist/config-map.yaml ships, redacted of anything that is a credential (those come from maximalist-secret.yaml and expand with ${VAR}):

tier: full

network:
geocoding: true # nominatim.openstreetmap.org: place names, trip names
map_tiles: true # server.arcgisonline.com: the trip fly-over, location cards

cache:
video_cache_max_size_gb: 5 # well inside the base's 20Gi cache PVC; a bigger
thumbnail_cache_max_size_mb: 3000 # library raises both, and the PVC first

advanced:
auth:
enabled: true
provider: oidc # Auth0, behind Traefik terminating TLS
public_url: "https://memories.example.com"
trusted_proxies:
- "10.42.0.0/16" # the cluster's pod CIDR, v4
- "fd42::/48" # and v6
issuer_url: "${OIDC_ISSUER_URL}"
client_id: "${OIDC_CLIENT_ID}"
client_secret: "${OIDC_CLIENT_SECRET}"
allowed_emails:
- "you@example.com"

server:
secure_cookies: true

editorial:
preparation:
caption_base_url: "http://captioner:8092/v1" # gpu-node-b, server-cuda

inference:
facts_base_url: "http://inference:8092" # gpu-node-a, ONNX Runtime on CUDA

llm:
provider: "openai-compatible"
base_url: "http://192.168.1.50:9999/v1" # oMLX on the Mac, LAN
model: "gemma-4-e4b-it-6bit"
api_key: "${LLM_API_KEY}"

ace_step:
enabled: true
mode: api
api_url: "http://acestep-api:8001" # an ACE-Step 1.5 API server, not in the overlay
api_key: "${ACE_STEP_API_KEY}"

automation:
enabled: true
daily_at: "09:00"

Music stems​

ACE-Step generates the full track in its separate deployment. Demucs in the inference service splits it into drums, bass, other and vocals for separate ducking under the clips. The existing advanced.inference.facts_base_url also selects /audio/stems. This setup needs ACE-Step and Demucs; a MusicGen server is optional and is not deployed by this repository.

The inference CUDA image includes the 80 MB htdemucs weights under /opt/immich-models/torch. The CPU inference image downloads them on first separation into /cache/torch; keep /cache on the model-cache PVC. Jobs run one at a time and their temporary audio files are removed after the response is sent. ACE-Step keeps its own image, models and Terraform deployment. On NVIDIA cards without BF16 support, its Oobleck VAE needs FP32 through CPU offload/reload: FP16 can produce NaNs and silent tracks. The ACE-Step image patch sets that precision in the VAE selector. Deploy the rebuilt image digest; this upstream selector does not read ACESTEP_DTYPE.

If the service fails, advanced.inference.fallback_to_local (default true) allows local Demucs. The app and render-worker image includes it on CPU; the Mac uses Metal. Local separation caches weights under ~/.cache/torch/hub/checkpoints. Set TORCH_HOME to a persistent cache directory to reuse them after a container restart. With no inference URL, local separation remains the default. An explicitly enabled MusicGen server retains priority for stems for existing installations.

The app image takes both Torch and TorchAudio from the CPU wheel index. The inference images pin matching CPU or CUDA 12.8 wheels. Every image checks the real Demucs model-loader import at build time, catching missing native libraries before release.

The two GPU nodes​

gpu-node-a has the newer card, here an NVIDIA T1000 (Turing, 8 GB), time-sliced with Immich's own machine-learning pod. It carries the app pod with the render worker as a loopback sidecar (NVENC h264/hevc, exec probes, below) and the inference service, which reads every picture once: the encoder, the context heads and the two detectors, on ONNX Runtime's CUDA provider. gpu-node-b has an older GTX 1070 (Pascal, 8 GB) and carries only the caption server: llama.cpp's CUDA build runs on Pascal, and the caption server is the one GPU workload that doesn't compete with the app for the busier card.

overlays/inference-cuda and overlays/captioner-cuda select any node with nvidia.com/gpu.present, so on a two-card cluster pin each to its node. The inference service belongs on the newer card; this setup was not tested with it on Pascal:

# a patch in overlays/maximalist, listed under patches:
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-memories-inference
namespace: immich-memories
spec:
template:
spec:
nodeSelector:
kubernetes.io/hostname: gpu-node-a # your newer card's node

Gotchas, with the exact strings to search for​

OIDC, unpinned public_url. Without auth.public_url the redirect_uri sent to the IdP is built from the in-cluster request and comes out http://, which every IdP refuses before the app ever sees the callback.

OIDC, untrusted proxy. With public_url set but no auth.trusted_proxies, X-Forwarded-Proto from Traefik is not trusted, so the callback still looks like plain HTTP and comes back:

400 {"detail":"Invalid callback origin"}

trusted_proxies is the pod CIDR your CNI hands out, IPv4 and IPv6.

enableServiceLinks left on. Kubernetes injects a set of env vars for every Service in the namespace by default, and this app's own IMMICH_MEMORIES_* prefix collides with its own Service names. A Service named immich-memories-render-worker injects IMMICH_MEMORIES_RENDER_WORKER_PORT=tcp://10.x.x.x:8093, which the worker's settings then read as its own port field and crash on (pydantic's validation error):

port
Input should be a valid integer, unable to parse string as an integer [type=int_parsing, input_value='tcp://…:8093', input_type=str]

Every manifest here sets enableServiceLinks: false for exactly this reason (#1608); copy it onto any pod spec you write yourself.

A tcpSocket/httpGet probe on a loopback-only worker. The kubelet dials the pod IP for both of those, never 127.0.0.1: a process inside the container makes no difference. The render worker binds loopback only, so either probe kind fails forever and the pod never goes Ready. The sidecar's startupProbe/readinessProbe run exec instead, inside the worker's own network namespace, where loopback is reachable.

A PVC that never binds. A storage class provisioning only static, pre-created PVs (no dynamic provisioner, volumeBindingMode: Immediate) leaves a fresh PVC Pending forever if no PV happens to match it. The caption weights are the case that bites here: they are pinned artifacts, cheap to refetch, so on a storage class that works this way, patch the caption server's models volume to an emptyDir in place of the immich-memories-caption-models claim overlays/captioner ships. The cache, output and models PVCs the app itself needs still want real storage.

tier: full in the file, nas in the pod. The base Deployment's IMMICH_MEMORIES_TIER=auto beats config.yaml, and auto stays on the plain NAS tier until it finds GPU picture reading. Nothing fails: films still render, from the rules alone, and the caption server and the reader simply never get a request. Search the log for Automatic selection tier: nas, and check with config show as above. The overlay sets the tier in the environment for this reason.

Which card gets what. llama.cpp's CUDA build runs on Pascal (sm_61); PyTorch's cu128 wheels no longer do. Put the caption server on the older card, and the inference service and the render worker on the newer one.

MTU on a multi-site cluster. A cluster spanning two sites over VXLAN loses bytes to the tunnel header; Cilium's default MTU assumes a 1500-byte path that isn't there. Symptoms are intermittent: large responses (a picture upload, a chat completion from the LAN LLM) stall or hang while small ones work. Set the CNI's MTU to match the tunnel's real ceiling, e.g. 1370 for a ~1420-byte path, rather than discovering it one timeout at a time.

The laptop / workstation (the Mac)​

Nothing above needs a second machine or a cluster; this profile runs the same app, the same config keys, entirely on one Mac, with two local servers instead of a cluster. lib mode is not in uv tool install or the all-mac extra: ACE-Step runs from a .venv-acestep beside a checkout (Install locally on a Mac). OIDC needs authlib, which all-mac and make dev-mac leave out; make dev installs every extra and builds the web client (it needs Node 22):

git clone https://github.com/sam-dumont/immich-video-memory-generator.git
cd immich-video-memory-generator
make dev
make install-acestep
uv run immich-memories ui

The Mac's config.yaml, annotated​

Every key below is valid on current main; nothing here is exotic or Tier-2-only by accident.

tier: full

network:
geocoding: true
map_tiles: true

cache:
video_cache_max_size_gb: 10
thumbnail_cache_max_size_mb: 10000

advanced:
auth:
enabled: true
provider: oidc # the same IdP as the cluster profile
issuer_url: "${OIDC_ISSUER_URL}"
client_id: "${OIDC_CLIENT_ID}"
client_secret: "${OIDC_CLIENT_SECRET}"

editorial:
preparation:
# mlxcel, serving the same SmolVLM2 alias as the llama.cpp recipe
caption_base_url: "http://localhost:8092/v1"

llm:
provider: "openai-compatible"
base_url: "http://localhost:9999/v1" # oMLX, also the cluster's reader over the LAN
model: "gemma-4-e4b-it-6bit"

ace_step:
enabled: true
mode: lib # a local library, not an API server
model_variant: "acestep-v15-xl-turbo" # the XL variant
lm_model_size: "4B"
use_lm: true

mode: lib needs Python 3.12 specifically (ACE-Step config reference); mode: api (the cluster profile's choice) has no such constraint, which is why the two profiles differ here. Running oidc against issuer_url alone, with no public_url, is fine on a single-user machine reached only at localhost: nothing forwards a proxied Host header for trusted_proxies to matter, so the origin check in Authentication has nothing to second-guess.

Terraform​

deploy/terraform/examples/maximalist is the Terraform form of the server profile: every feature above as a module variable, defaulting to the minimal path (off) until set, except the tier, automation and the inference service's URL, which go through env. The module does not deploy the inference service: apply deploy/kubernetes/overlays/inference-cuda beside it. See deploy/terraform/README.md.