Skip to main content

Requirements and tiers

The default install is one container on the box that already runs Immich, and it makes the whole film there. A GPU or a text model makes it better: what each one adds, feature by feature. This page is what each setup needs, and how the app picks between them.

Hardware​

For this app's container, on top of what Immich itself uses:

MinimumRecommended
RAM4 GB free for the container (the compose file's limit)8 GB, for 4K output or a render running beside Immich's own jobs
CPU2 cores, x86-64 or ARM644 cores, x86-64 with AVX
Disk25 GB on the config volume, plus the 2.4 GB image and your filmsthe config volume on an SSD
OSLinux with Docker Engine and Compose v2same; Docker Desktop on a Mac or Windows works too
Immichv2 or v3, and an API keysame

What the minimum costs you:

  • Two cores make the render the long part of every run. The editor banks what it reads, but not the encode: a second cut of the same month reads nothing again and still encodes the whole film.
  • No AVX (Intel Celeron J4125 and friends) means the CPU fallback draws the titles instead of the animated title kernels: CPUs without AVX.
  • ARM64 gets no hardware encoder: the VA-API drivers ship in the amd64 image only.
  • Less memory means fewer photos rendered at once. The app prepares one source per 2 GB it may use: the container's memory limit when Compose sets one (the shipped file sets 4 GB), otherwise the machine's RAM. 2 or 3 GB renders one photo at a time, 4 GB and up renders two. immich-memories preflight prints what it picked, for example Photo preparation: 1 at a time (2.0 GB available, container limit). Setting advanced.analysis.source_prepare_workers to a number (1 to 4) overrides it. The same memory figure caps the threads of each clip decode in the render (one per 2 GB, up to 4): FFmpeg's own default of one per core cost 1.2 GB per 4K decode on an 18-core Mac. A box with no hardware HEVC encoder encodes in libx265, which holds about 52 MB per frame it looks ahead at 4K. Above 1080p the app lets it look 5 frames ahead up to 3 GB, 10 at 4 or 5 GB, and x265's default (20 at the medium preset) from 6 GB. The files come out a few percent smaller at a slightly lower quality: at 1080p with a lookahead of 10, 3% smaller and 0.03 dB lower. preflight shows the choice on its Memory line. 1080p output keeps the default everywhere. Below 3 GB there is no room for a 4K software HEVC film at all, so when the film's resolution is auto and the box has no hardware HEVC encoder, a 4K film renders at 1080p instead, in the same orientation. preflight and the run log say so. A resolution you set yourself, in the config or with --resolution, is kept: a 4K set that way below 3 GB gets a warning that the render may run out of memory, and the fix is auto or 1080p.

The 25 GB covers the caches at their default budgets (10 GB of Immich previews, 10 GB of downloaded video kept 7 days) with room for the store to grow. The models are about 140 MB. The one file worth backing up is the store, store.db, where every fact the editor read is banked: What to keep.

Which hosts it has run on, and when, is in Supported and tested below. Timings are on Measured.

Supported and tested​

Tested means run end to end, with the date and the commit or release it ran on: check the date against your version. Supported means the code path exists and worked on an earlier release, but has not been checked since: it probably works, and a report is welcome if it doesn't. Untested means nobody has run it; it may work.

The last release on PyPI is 0.103.0, from 2026-09-17. Rows tested after that date ran on main and the Docker image built from it, not on a pip install.

AreaWhatStateEvidence
SetupPlain NAS (nas tier)TestedEvery pull request cuts a month on a real Immich (v2 and v3); a cut, an edit and a render in the browser, 2026-09-27; 28 films on the maintainer's library (years, months, trips, seasons, people, special days), 2026-09-27, commit 9eb16812
SetupGPU and model (full tier)TestedA July film on the maintainer's library, on main, Apple Silicon, 2026-09-28; 28 films on the maintainer's library (years, months, trips, seasons, people, special days), 2026-09-27, commit 9eb16812
SetupGPU (gpu tier)Tested28 films on the maintainer's library (years, months, trips, seasons, people, special days), 2026-09-27, commit 9eb16812
Immichv2.7.5 and v3.2.2TestedChecked on every pull request
Immich3.1.0TestedThe maintainer's library, 2026-09-28
ImmichOther 2.x and 3.x releasesSupportedThe version is detected at runtime; only the three above are exercised
DatabaseSQLite (the default) and PostgreSQL 16TestedBoth on every pull request that touches the store, since 2026-09-28
Python3.11, 3.12, 3.13TestedEvery pull request, Linux and macOS
PlatformApple Silicon, from sourceTestedThe full tier film above, 2026-09-28
PlatformDocker on x86Tested, deployment onlyEvery image change starts the compose file, upgrades, backs up and restores the store; no film renders in that check
PlatformDocker on arm64UntestedThe image builds; it has not been run
PlatformSynology DS423+ (no AVX)SupportedA one-month film on 2026-09-17, release 0.102.0
PlatformKubernetes manifestsSupportedFilms on the maintainer's cluster, 2026-09-13 to 17
PlatformTerraform moduleUntested as shippedAn example module: adapt it to your cluster
GPUNVIDIA inference service and NVENC encodingSupported2026-09-17, release 0.102.0, on a T1000
GPUIntel VA-API and Quick SyncSupported2026-09-11, on the DS423+
GPUAMD VA-APIUntestedThe drivers ship in the image
Render workerThe service's own test suiteTestedEvery pull request that touches it; no dated deployment on a real GPU box
ReaderLocal: oMLX with Gemma 4 E4B (6-bit)Tested2026-09-27, commit 9eb16812, the full films above
ReaderLocal: llama.cpp, OllamaSupportedFilms on earlier releases
ReaderLocal: vLLM, mlx-vlm served directlyUntested
ReaderHosted: z.ai (glm-5.3-flash) and OpenAI (gpt-5.6-luna)SupportedLast run 2026-09-17; re-test: #1513
ReaderHosted: Melious (DeepSeek, deepseek-v4.1-flash)Supported; schema fallback, or advanced.llm.structured_output: falseLast run 2026-09-15; re-test: #1513
ReaderHosted: Anthropic's own APIUntestedThe same code path only ran through z.ai's Anthropic-compatible route
ReaderHosted: Melious gemma-4-31bNot supportedIts API refused every image (HTTP 400), 2026-09-15
CaptionsSmolVLM2 500M, on a MacTested2026-09-27, commit 9eb16812, the gpu and full films above
CaptionsSmolVLM2 500M, on the CUDA inference serviceSupported2026-09-17, release 0.102.0
CaptionsSmolVLM2 under llama.cppUntested
LayaThe family-viewing pre-screen, on a MacTested2026-09-27, commit 9eb16812, the gpu and full films above
LayaThe family-viewing pre-screen, on CUDASupportedFilms on earlier releases

The three tiers​

tier: auto, the default, picks one tier for preparation and selection alike:

TierPicked whenAlso needsWhat runsFamily-viewing check
nasNo GPU inference is found (the default)models fetchImmich metadata, and the DINOv2 encoder with eight heads and two detectors on the CPURules and the detectors
gpuA GPU inference runtime is found: the inference service reporting CUDA, a local CUDA runtime, or a Mac's Metal GPUA caption server and the Laya checkpointNAS, plus captions and Laya for the pictures in the cut and the candidates to replace themLaya can add holds; it never lifts one
fullGPU, plus a text model whose llm.model and endpoint (llm.base_url, or a hosted llm.provider) are both setA text model with a 32k contextGPU, plus the text model's account of the period, its polish of the draft, the title and the music moodSame as GPU. The text model never decides what is shareable

auto looks only at the GPU runtime and those llm keys, not at the caption server or the Laya checkpoint. immich-memories preflight checks the caption server, and models fetch also downloads the Laya checkpoint once the tier is gpu or full.

A render worker or hardware encoding moves or speeds up the encode. Neither changes the tier: a GPU that encodes video is not a GPU that runs the models.

Which tier you get​

  • A text model without the GPU tier still writes titles and picks the music mood. Selection stays on nas, and the log says which service is missing.
  • On a Mac the mac extra's Metal bindings find the GPU, so an all-mac install reaches gpu and full; the caption server and the reader run as their own processes.
  • IMMICH_MEMORIES_TIER beats tier: in config.yaml. The compose file and the Kubernetes manifests set it to auto: remove it if you want the file to decide.
  • nas, gpu and full can be set explicitly, for side-by-side comparisons. They don't install a model or start a service. immich-memories preflight checks what the resolved tier needs.
  • A cut reads the cheap facts for every picture it can reach. Captions and Live Photo checks wait for the pictures it selects and their replacement candidates. immich-memories prepare reads a whole period ahead of time when you ask for it.
  • Everything is banked per picture and per producer, so changing the tier erases nothing, and a NAS library can add captions later.