Skip to main content

Basic with a local text model

This reference adds Ollama 0.35.1, model gemma4:e4b-it-q4_K_M, to the prebuilt Basic app. It supplies a reader: a text model (LLM) configured under advanced.llm. It can write titles and music mood. It does not add GPU selection, Full or vision captions. Full additionally needs GPU inference, the caption service and Laya.

The published Docker image has no llama-server: Docker and Kubernetes require an external text-model server. Native Linux/macOS can instead run the app-owned llama.cpp reader, which is a different model/runtime recipe. External services own their memory and shutdown.

This exact model has recorded conformance results: validated for text use, with motion interpretation as a known weak spot, not a strength. Other providers and their failures are tracked separately.

Prepare the two services​

Start from the Basic prebuilt installation. Use an operator-managed Ollama host reachable only from the app's private network; its installation, RAM/VRAM and outbound firewall belong to that host. On the model host, verify the version and acquire the exact model:

ollama --version
ollama pull gemma4:e4b-it-q4_K_M
ollama show gemma4:e4b-it-q4_K_M

Keep the displayed model identity/digest alongside your config: an Ollama tag alone can move. Ollama's native endpoint is port 11434, with no app-supplied authentication in this recipe. Keep it on a restricted LAN/VLAN or behind your authenticated proxy; never publish it to the internet. Its model files stay in Ollama's model directory on that host, separate from app data.

Save this as config.yaml beside the app Compose file. 192.168.1.50 is an example; replace it with the actual private model-host IP. Retain the Basic app's .env for Immich and login secrets.

tier: basic
upload:
enabled: false
network:
geocoding: false
map_tiles: false
advanced:
llm:
enabled: true
provider: ollama
base_url: http://192.168.1.50:11434
model: gemma4:e4b-it-q4_K_M
reader_concurrency: 1
extra_params:
think: false
options:
num_ctx: 32768
ace_step:
enabled: false
musicgen:
enabled: false
notifications:
urls: []

Add a read-only file mount under the app service's volumes:

- ./config.yaml:/home/immich/.immich-memories/config.yaml:ro

The app owns advanced.llm; Ollama owns model acquisition, port/listen policy and resident memory. File/environment values override saved Settings. Do not also enable a hosted provider. A dotted LAN hostname receives the app's “hosted” request defaults; that classification does not establish where traffic goes. This private-IP recipe avoids it. For dotted names, see request concurrency and schema overrides.

On the app host, recreate, prepare and probe:

docker compose up -d
docker compose exec immich-memories immich-memories models fetch
docker compose exec immich-memories immich-memories preflight
docker compose exec immich-memories immich-memories capabilities

Expect Basic selection and an enabled Ollama reader. A model name alone does not establish readiness: fix failed reader connectivity first. Keep bundled music, local output and the 20-50-item trial. Reader failure falls back to rules and default wording rather than stopping the run; preflight reports it (LLM WARNING Cannot connect, exit 0), and the run log carries one line: Reader unreachable at <host> (<error>); falling back to rules and default wording. Check for that line before counting a film as reader-assisted.

Network phases and verification​

PhaseExpected recipients and data
Artifact/model preparationApp registry and delivery hosts for image layers; GitHub/CDNs for Basic pins, WordNet source; Ollama registry/storage hosts for model weights
Startup and warm-upApp → Immich for API/auth probes; app → private Ollama for model/request checks; model host loads its prepared weights
Film generationApp → Immich for metadata/previews/originals; app → Ollama for text requests; local film output. No vision frames are opted into the text reader

Maps/geocoding, hosted endpoints, notifications and generated music stay off. This avoids ACE-Step/Demucs first-use downloads, but all participating services still need an outbound boundary. The request inventory lists each feature's recipients.

To close the network down to just this pair, finish preparation first, then use host/router firewall rules permitting app → Immich IP/port and app → Ollama IP/11434. Deny other app egress on IPv4 and IPv6. Deny the Ollama host/process public egress too, allowing only replies to the app and any strictly internal resolver needed. For containers, enforce forwarded traffic, not just the host OUTPUT chain; for Kubernetes apply destination policies to both pods and remove other additive allow rules. The offline guide details those enforcement limits. Verify from each namespace: a permitted local request succeeds, a public HTTPS request is denied, then run a bounded film and check the destinations logged during startup and generation.

Kubernetes form​

Same recipe, three differences: the config is a ConfigMap, Ollama runs as a pod in the namespace, and NetworkPolicies replace the firewall. Start from the Kubernetes installation; kubectl apply -k your own root, never base files singly.

Hardware. Check the model fits your node's free VRAM before requesting a GPU; without one, Ollama runs on CPU. Titles and mood are one or a few calls per film, so plan for minutes, not seconds, per call without a GPU. Measured numbers for this setup are on the measurements page.

Secret. Put IMMICH_MEMORIES_SECRET_KEY in immich-memories-secrets next to IMMICH_URL and IMMICH_API_KEY (the Quick start's secret.yaml.example has it). Without it, every config show prints "secrets cannot be saved".

Config. Save the config.yaml from above, with one change: base_url: http://ollama:11434. Use the short Service name with no dots: a dotted name gets the app's hosted request defaults, and ollama resolves inside the namespace. The kustomization generates the ConfigMap and mounts it with subPath, so the rest of ~/.immich-memories stays writable. The mount goes on the app and its init container:

# kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: <your namespace>
resources:
- <path to deploy/kubernetes/base>
- ollama.yaml
# - networkpolicy-offline.yaml # uncomment after the model is pulled, see below
configMapGenerator:
- name: immich-memories-config
files:
- config.yaml
generatorOptions:
disableNameSuffixHash: true
patches:
- target: {kind: Deployment, name: immich-memories}
patch: |-
- op: add
path: /spec/template/spec/volumes/-
value: {name: app-config, configMap: {name: immich-memories-config}}
- op: add
path: /spec/template/spec/containers/0/volumeMounts/-
value: {name: app-config, mountPath: /home/immich/.immich-memories/config.yaml, subPath: config.yaml, readOnly: true}
- op: add
path: /spec/template/spec/initContainers/0/volumeMounts/-
value: {name: app-config, mountPath: /home/immich/.immich-memories/config.yaml, subPath: config.yaml, readOnly: true}

Ollama. A Deployment (Recreate, one replica), a Service named ollama and a 20Gi claim for the models. No GPU request; add one only after checking the card holds the model.

# ollama.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: ollama-models
labels: {app.kubernetes.io/name: ollama}
spec:
accessModes: [ReadWriteOnce]
resources: {requests: {storage: 20Gi}}
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: ollama
labels: {app.kubernetes.io/name: ollama}
spec:
replicas: 1
strategy: {type: Recreate}
selector: {matchLabels: {app.kubernetes.io/name: ollama}}
template:
metadata:
labels: {app.kubernetes.io/name: ollama}
spec:
enableServiceLinks: false
automountServiceAccountToken: false
containers:
- name: ollama
image: ollama/ollama:0.35.1
env:
- {name: OLLAMA_HOST, value: "0.0.0.0:11434"}
ports: [{name: http, containerPort: 11434}]
resources:
requests: {cpu: "2", memory: 8Gi}
limits: {cpu: "6", memory: 14Gi}
readinessProbe:
httpGet: {path: /, port: http}
periodSeconds: 10
volumeMounts:
- {name: models, mountPath: /root/.ollama}
volumes:
- name: models
persistentVolumeClaim: {claimName: ollama-models}
---
apiVersion: v1
kind: Service
metadata:
name: ollama
labels: {app.kubernetes.io/name: ollama}
spec:
selector: {app.kubernetes.io/name: ollama}
ports: [{name: http, port: 11434, targetPort: http}]

Apply, then pull the model once while the Ollama pod can still reach the registry (the policy below blocks that), and fetch the app's pins:

NS=<your namespace>
kubectl apply -k <your root>
kubectl exec -n "$NS" deploy/ollama -- ollama pull gemma4:e4b-it-q4_K_M
kubectl exec -n "$NS" deploy/immich-memories -- immich-memories preflight -v

The pull takes a few minutes cold, depending on your cluster's bandwidth. preflight should print LLM OK Connected (ollama, 1 models).

NetworkPolicies, after the pull. Uncomment the file in the kustomization and apply again. Two policies, one per pod. The app policy is the offline example with one more rule, and it replaces the base policy under the same name. Immich is selected by pod, not by IP (see that page for why). The Ollama pod may only answer the app and resolve DNS.

# networkpolicy-offline.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: immich-memories
spec:
podSelector:
matchLabels: {app.kubernetes.io/name: immich-memories}
policyTypes: [Ingress, Egress]
ingress:
- ports: [{port: 8080, protocol: TCP}]
egress:
- to:
- podSelector:
matchLabels:
app.kubernetes.io/name: immich-memories
app.kubernetes.io/component: web-ui
ports: [{port: 8080, protocol: TCP}]
- to:
- namespaceSelector:
matchLabels: {kubernetes.io/metadata.name: kube-system}
podSelector:
matchLabels: {k8s-app: kube-dns}
ports: [{port: 53, protocol: UDP}, {port: 53, protocol: TCP}]
- to: # Immich in this cluster: its namespace and pod labels, not its LoadBalancer IP
- namespaceSelector:
matchLabels: {kubernetes.io/metadata.name: <immich namespace>}
podSelector:
matchLabels: {<immich pod label>: <value>}
ports: [{port: 2283, protocol: TCP}]
- to:
- podSelector:
matchLabels: {app.kubernetes.io/name: ollama}
ports: [{port: 11434, protocol: TCP}]
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: ollama
spec:
podSelector:
matchLabels: {app.kubernetes.io/name: ollama}
policyTypes: [Ingress, Egress]
ingress:
- from:
- podSelector:
matchLabels: {app.kubernetes.io/name: immich-memories}
ports: [{port: 11434, protocol: TCP}]
egress:
- to:
- namespaceSelector:
matchLabels: {kubernetes.io/metadata.name: kube-system}
podSelector:
matchLabels: {k8s-app: kube-dns}
ports: [{port: 53, protocol: UDP}, {port: 53, protocol: TCP}]

Applying the root again restores the base's broad egress policy; keep this file in the root so it wins. Verify the policy from inside the pods: app to Immich and to ollama:11434 should succeed, app to any public host should time out, and the app's logs should name no host besides Immich and ollama:11434.

Stopped provider. kubectl scale deploy/ollama -n "$NS" --replicas=0 makes preflight print LLM WARNING Cannot connect (exit 0). A film still renders, with the bundled track and the default mood and no hosted fallback. Scale back to 1 and preflight returns to LLM OK once Ollama is ready again.

Kubernetes services​

Use the same candidate/release pins for project images. VERSION below means that exact app image tag, never latest; the published bundle substitutes it. CPU app/inference images target amd64 and arm64; CUDA inference targets amd64 only. Inspect the published manifest before selecting an architecture. Third-party caption/reader image support is independent.

ServiceImage / endpoint / authCPU and RAM requests → limitsGPU and storage / startup
App, one replicaProject app :VERSION; 8080; Basic/OIDC before LAN exposure1 CPU/2Gi → 4 CPU/8Gi; fetch init 250m/512Mi → 2 CPU/2GiNo GPU in Basic or remote-inference wrapper. Data PVC contains SQLite, Settings/credentials and caches; /models pins; /app/output films. Init fetches pins
CUDA inferenceProject /inference:VERSION-cuda; 8092; no key in shipped overlay, restrict network500m/1Gi → 4 CPU/4GiOne GPU allocation; /cache PVC for weights/HF/JIT; first use can download and warm models; idle unload does not remove files
Generated CUDA captionerghcr.io/ggml-org/llama.cpp:server-cuda-b10920; 8092 /v1, alias smolvlm2-500m-base-public; network isolation500m/1Gi → 4 CPU/3Gi; fetch init 100m/64Mi → 1 CPU/256MiTime-sliced, no reservation in the generated GPU wrapper; /models PVC contains pinned SmolVLM2 Q8_0 and projector; init verifies/downloads them
Independently managed CUDA captioner overlaySame caption server/alias and resource limitsSame as aboveNo GPU reservation in this overlay: operator must provide supported device access/sharing or request a separate card. It still consumes VRAM
External Full readerOperator's tested server/model; endpoint and API key as configuredOperator budget; no common measured minimumIts own model cache, VRAM and startup/download policy. App cannot unload external weights
Optional render workerWorker image/config; 8093; bearer tokenOperator-selected worker/sidecar budgetEncoding GPU is separate from selection capability; scratch and model paths belong to worker; media arrives over authenticated HTTP

The generated GPU wrapper reserves one GPU allocation, for the inference service. The captioner reserves none and relies on time-slicing; the wrapper does not configure sharing, so a single exclusive card is not guaranteed to be enough for both. VRAM depends on loaded models and concurrency; a missing reservation does not mean zero consumption.

Only the app and its init/maintenance jobs need the app state volumes. Inference, caption and external reader use their own caches; HTTP endpoints do not require sharing SQLite or app paths. Keep SQLite on local/block storage. PostgreSQL removes that file constraint but does not make multiple UI replicas safe. See mounts, backups and probes. Existing services can be reused once preflight confirms the expected facts contract, caption alias and reader configuration.