Skip to main content

Distributed services on Kubernetes

Use separate services when you need independent GPU placement or already run a reader and music server. Start with the base Kubernetes install. For one GPU shared between preparation and rendering, the one GPU service is simpler.

Deployment explorerDistributed Kubernetes services
Existing Immich server

Your Immich

Your photo and video library
App pod · one replica

Immich Memories

Data/cache · output · model PVCs
Pods + supplied reader endpoint

Preparation services

Service model caches
1

Metadata + previewsPreparation starts with your library; originals remain in Immich.

2

Preview tiles + framesClassification and captions go to separate services. The text reader receives annotations through its separately configured endpoint.

Service boundaries, not node placement. Reserve GPUs explicitly; a namespace is not an authentication boundary. Reader and music endpoints are configured separately. Select a service for details.

How I run it describes the maintainer's own cluster built from these pieces, why each one is there, and where it departs from this overlay.

What runs where​

The shipped deploy/kubernetes/overlays/maximalist example composes these pieces. Use it as a starting point and keep the services you need.

ComponentDeployed by the overlayNeeds
App and persistent storeYes, one app replicaImmich access and the base PVCs
Render workerYes, sidecar in the app podNVIDIA GPU allocation; shared bearer token
Classifiers and Demucs stemsYes, separate inference podNVIDIA GPU allocation; model cache
CaptionsYes, separate caption podNVIDIA runtime; reviewed GPU-sharing settings
Text readerNoExisting compatible endpoint with 32k context
ACE-Step musicNoExisting ACE-Step API server, if music generation is enabled
HTTPS proxy and identity providerNoIngress/proxy and OIDC registration, if keeping the example's login settings

Laya remains in the app. The reader receives selection text; inference and captions receive picture previews. The render worker downloads originals from Immich with the API key supplied by the app. Privacy lists the data sent to each service.

Before applying​

Keep the storage, namespace and NVIDIA prerequisites from the base install. The render sidecar and inference pod each claim one advertised GPU allocation. Putting both on one physical card requires device-plugin sharing configured by the cluster operator; the overlay does not set that up. Choose node selectors that match your available allocations.

The caption overlay intentionally has no nvidia.com/gpu request. It can share devices exposed by the NVIDIA runtime, but this provides no scheduler reservation or memory guarantee. For a dedicated card, add the explicit GPU request and set its node selector. Confirm driver/runtime support and memory headroom for every service.

Use a release deployment bundle or a checkout of the release you will run. Keep the app and render worker tags equal. Inspect image pins in the base and each add-on kustomization; the caption image uses a separate llama.cpp build. Caption image pinning explains that serving contract.

Configure your endpoints and access​

From the deployment bundle or release checkout:

cd deploy/kubernetes
cp base/secret.yaml.example base/secret.yaml
cp overlays/render-sidecar/render-worker-secret.yaml.example overlays/render-sidecar/render-worker-secret.yaml
cp overlays/maximalist/maximalist-secret.yaml.example overlays/maximalist/maximalist-secret.yaml
openssl rand -hex 32

Edit those three Secrets and overlays/maximalist/config-map.yaml before applying:

File or settingReplace or confirm
Base SecretYour Immich URL and API key
Render-worker SecretSame Immich URL and the generated token
Extra SecretYour OIDC credentials and any reader/music API keys
advanced.llmReader URL and the exact model name served there
advanced.ace_stepMusic API URL, or enabled: false to use bundled/chosen music
advanced.auth, advanced.serverPublic HTTPS URL, allowed accounts, trusted proxy addresses and secure cookies
network and cache limitsOutside map calls you permit and limits that fit the PVC

The example enables OIDC and assumes HTTPS. Configure your own authentication and proxy, and register its callback/logout URLs with the identity provider. The overlay does not create an Ingress. For a private port-forward install, replace the OIDC settings with the access method you chose; keep the app private until that method works.

The additional NetworkPolicy permits reader port 9999 and music port 8001. Change those ports with your endpoint URLs. These are port rules, not destination allow-lists. The inference and caption listeners have no built-in authentication; keep them private and restrict their clients in your network policy. A ClusterIP does not authenticate requests.

Render, apply and verify​

kubectl kustomize overlays/maximalist
kubectl apply -k overlays/maximalist
kubectl rollout status -n immich-memories deploy/immich-memories
kubectl rollout status -n immich-memories deploy/immich-memories-inference
kubectl rollout status -n immich-memories deploy/immich-memories-captioner
kubectl get pods -n immich-memories -o wide
kubectl exec -n immich-memories deploy/immich-memories -c immich-memories -- immich-memories models fetch
kubectl exec -n immich-memories deploy/immich-memories -c immich-memories -- immich-memories config show tier
kubectl exec -n immich-memories deploy/immich-memories -c immich-memories -- immich-memories preflight -v

The overlay explicitly sets IMMICH_MEMORIES_TIER=full: the environment wins over the config file. Seeing Full in config show confirms the setting; preflight checks its required services and models. Run preparation on a small date window, review a first film and confirm its reported render path before enabling daily automation.

If a pod stays Pending, check GPU allocations, node selectors and PVC events. If preflight fails, check the relevant endpoint, model name, credentials and NetworkPolicy before changing tiers. Keep the worker's loopback exec probes; kubelet HTTP probes cannot reach its loopback listener. Service diagnostics has the logs and health routes.

Adapt the example​

Remove unwanted services from your own composition rather than copying every option. Applying sibling app overlays one after another replaces earlier patches; list the app components in one kustomization instead. Preserve enableServiceLinks: false and one app replica.

For PostgreSQL, add the database Secret and egress port to that composition; the overlay does not deploy PostgreSQL. Database setup and overlay composition cover the details.

Terraform users can start from deploy/terraform/examples/maximalist. That example creates the app, render sidecar and captions, but not inference: deploy inference separately in the same namespace and match the configured URL. Reader, music and identity services remain external. Terraform inputs describe those switches.