Requirements and tiers
The default install is one container on the box that already runs Immich, and it makes the whole film there. A GPU or a text model makes it better: what each one adds, feature by feature. This page is what each setup needs, and how the app picks between them.
Hardware
For this app's container, on top of what Immich itself uses:
| Minimum | Recommended | |
|---|---|---|
| RAM | 4 GB free for the container (the compose file's limit) | 8 GB, for 4K output or a render running beside Immich's own jobs |
| CPU | 2 cores, x86-64 or ARM64 | 4 cores, x86-64 with AVX |
| Disk | 25 GB on the config volume, plus the 2.4 GB image and your films | the config volume on an SSD |
| OS | Linux with Docker Engine and Compose v2 | same; Docker Desktop on a Mac or Windows works too |
| Immich | v2 or v3, and an API key | same |
What the minimum costs you:
- Two cores make the render the long part of every run. The editor banks what it reads, but not the encode: a second cut of the same month reads nothing again and still encodes the whole film.
- No AVX (Intel Celeron J4125 and friends) means the CPU fallback draws the titles instead of the animated title kernels: CPUs without AVX.
- ARM64 gets no hardware encoder: the VA-API drivers ship in the amd64 image only.
- Less memory means fewer photos rendered at once. The app prepares one source per 2 GB it
may use: the container's memory limit when Compose sets one (the shipped file sets 4 GB),
otherwise the machine's RAM. 2 or 3 GB renders one photo at a time, 4 GB and up renders two.
immich-memories preflightprints what it picked, for examplePhoto preparation: 1 at a time (2.0 GB available, container limit). Settingadvanced.analysis.source_prepare_workersto a number (1 to 4) overrides it. The same memory figure caps the threads of each clip decode in the render (one per 2 GB, up to 4): FFmpeg's own default of one per core cost 1.2 GB per 4K decode on an 18-core Mac. A box with no hardware HEVC encoder encodes in libx265, which holds about 52 MB per frame it looks ahead at 4K. Above 1080p the app lets it look 5 frames ahead up to 3 GB, 10 at 4 or 5 GB, and x265's default (20 at themediumpreset) from 6 GB. The files come out a few percent smaller at a slightly lower quality: at 1080p with a lookahead of 10, 3% smaller and 0.03 dB lower.preflightshows the choice on its Memory line. 1080p output keeps the default everywhere. Below 3 GB there is no room for a 4K software HEVC film at all, so when the film's resolution isautoand the box has no hardware HEVC encoder, a 4K film renders at 1080p instead, in the same orientation.preflightand the run log say so. A resolution you set yourself, in the config or with--resolution, is kept: a 4K set that way below 3 GB gets a warning that the render may run out of memory, and the fix isautoor1080p.
The 25 GB covers the caches at their default budgets (10 GB of Immich previews, 10 GB of downloaded
video kept 7 days) with room for the store to grow. The models are about 140 MB. The one file worth backing
up is the store, store.db, where every fact the editor read is banked:
What to keep.
Which hosts it has run on, and when, is in Supported and tested below. Timings are on Measured.
Supported and tested
Tested means run end to end, with the date and the commit or release it ran on: check the date against your version. Supported means the code path exists and worked on an earlier release, but has not been checked since: it probably works, and a report is welcome if it doesn't. Untested means nobody has run it; it may work.
The last release on PyPI is 0.103.0, from 2026-09-17. Rows tested after that date ran on main
and the Docker image built from it, not on a pip install.
| Area | What | State | Evidence |
|---|---|---|---|
| Setup | Plain NAS (nas tier) | Tested | Every pull request cuts a month on a real Immich (v2 and v3); a cut, an edit and a render in the browser, 2026-09-27; 28 films on the maintainer's library (years, months, trips, seasons, people, special days), 2026-09-27, commit 9eb16812 |
| Setup | GPU and model (full tier) | Tested | A July film on the maintainer's library, on main, Apple Silicon, 2026-09-28; 28 films on the maintainer's library (years, months, trips, seasons, people, special days), 2026-09-27, commit 9eb16812 |
| Setup | GPU (gpu tier) | Tested | 28 films on the maintainer's library (years, months, trips, seasons, people, special days), 2026-09-27, commit 9eb16812 |
| Immich | v2.7.5 and v3.2.2 | Tested | Checked on every pull request |
| Immich | 3.1.0 | Tested | The maintainer's library, 2026-09-28 |
| Immich | Other 2.x and 3.x releases | Supported | The version is detected at runtime; only the three above are exercised |
| Database | SQLite (the default) and PostgreSQL 16 | Tested | Both on every pull request that touches the store, since 2026-09-28 |
| Python | 3.11, 3.12, 3.13 | Tested | Every pull request, Linux and macOS |
| Platform | Apple Silicon, from source | Tested | The full tier film above, 2026-09-28 |
| Platform | Docker on x86 | Tested, deployment only | Every image change starts the compose file, upgrades, backs up and restores the store; no film renders in that check |
| Platform | Docker on arm64 | Untested | The image builds; it has not been run |
| Platform | Synology DS423+ (no AVX) | Supported | A one-month film on 2026-09-17, release 0.102.0 |
| Platform | Kubernetes manifests | Supported | Films on the maintainer's cluster, 2026-09-13 to 17 |
| Platform | Terraform module | Untested as shipped | An example module: adapt it to your cluster |
| GPU | NVIDIA inference service and NVENC encoding | Supported | 2026-09-17, release 0.102.0, on a T1000 |
| GPU | Intel VA-API and Quick Sync | Supported | 2026-09-11, on the DS423+ |
| GPU | AMD VA-API | Untested | The drivers ship in the image |
| Render worker | The service's own test suite | Tested | Every pull request that touches it; no dated deployment on a real GPU box |
| Reader | Local: oMLX with Gemma 4 E4B (6-bit) | Tested | 2026-09-27, commit 9eb16812, the full films above |
| Reader | Local: llama.cpp, Ollama | Supported | Films on earlier releases |
| Reader | Local: vLLM, mlx-vlm served directly | Untested | |
| Reader | Hosted: z.ai (glm-5.3-flash) and OpenAI (gpt-5.6-luna) | Supported | Last run 2026-09-17; re-test: #1513 |
| Reader | Hosted: Melious (DeepSeek, deepseek-v4.1-flash) | Supported; schema fallback, or advanced.llm.structured_output: false | Last run 2026-09-15; re-test: #1513 |
| Reader | Hosted: Anthropic's own API | Untested | The same code path only ran through z.ai's Anthropic-compatible route |
| Reader | Hosted: Melious gemma-4-31b | Not supported | Its API refused every image (HTTP 400), 2026-09-15 |
| Captions | SmolVLM2 500M, on a Mac | Tested | 2026-09-27, commit 9eb16812, the gpu and full films above |
| Captions | SmolVLM2 500M, on the CUDA inference service | Supported | 2026-09-17, release 0.102.0 |
| Captions | SmolVLM2 under llama.cpp | Untested | |
| Laya | The family-viewing pre-screen, on a Mac | Tested | 2026-09-27, commit 9eb16812, the gpu and full films above |
| Laya | The family-viewing pre-screen, on CUDA | Supported | Films on earlier releases |
The three tiers
tier: auto, the default, picks one tier for preparation and selection alike:
| Tier | Picked when | Also needs | What runs | Family-viewing check |
|---|---|---|---|---|
nas | No GPU inference is found (the default) | models fetch | Immich metadata, and the DINOv2 encoder with eight heads and two detectors on the CPU | Rules and the detectors |
gpu | A GPU inference runtime is found: the inference service reporting CUDA, a local CUDA runtime, or a Mac's Metal GPU | A caption server and the Laya checkpoint | NAS, plus captions and Laya for the pictures in the cut and the candidates to replace them | Laya can add holds; it never lifts one |
full | GPU, plus a text model whose llm.model and endpoint (llm.base_url, or a hosted llm.provider) are both set | A text model with a 32k context | GPU, plus the text model's account of the period, its polish of the draft, the title and the music mood | Same as GPU. The text model never decides what is shareable |
auto looks only at the GPU runtime and those llm keys, not at the caption server or the
Laya checkpoint. immich-memories preflight checks the caption server, and models fetch also
downloads the Laya checkpoint once the tier is gpu or full.
A render worker or hardware encoding moves or speeds up the encode. Neither changes the tier: a GPU that encodes video is not a GPU that runs the models.
Which tier you get
- A text model without the GPU tier still writes titles and picks the music mood. Selection stays on
nas, and the log says which service is missing. - On a Mac the
macextra's Metal bindings find the GPU, so anall-macinstall reachesgpuandfull; the caption server and the reader run as their own processes. IMMICH_MEMORIES_TIERbeatstier:inconfig.yaml. The compose file and the Kubernetes manifests set it toauto: remove it if you want the file to decide.nas,gpuandfullcan be set explicitly, for side-by-side comparisons. They don't install a model or start a service.immich-memories preflightchecks what the resolved tier needs.- A cut reads the cheap facts for every picture it can reach. Captions and Live Photo checks wait for
the pictures it selects and their replacement candidates.
immich-memories preparereads a whole period ahead of time when you ask for it. - Everything is banked per picture and per producer, so changing the tier erases nothing, and a NAS library can add captions later.