Choose your setup
The Basic tier makes a complete film: stories, favourites, people, trips, time order, titles and bundled music. It runs on a NAS, a Mac or a cluster without GPU inference. Start there. Add services for a change you want to see in the film.
Setup files for docs build v1.0.0-rc.1. The app image and release assets use this same version.
Create a key with the ten required read permissions. Add the five upload permissions only to send films back; asset.delete is optional. Avoid the All preset.
Can I run this? Check the version and topology matrix.
Not yet tested as an end-to-end generated installation on Linux Docker Compose. Files are checked with Compose/Kustomize and the form is checked in a browser. Those checks do not run this installation. Earlier NAS, Mac and GPU Kubernetes checks are recorded in the measured results. Tried it? Report your platform, release and preflight result.
Your choices produce files in this browser. Nothing is sent to an Immich or model server.
Enter your Immich API key.
For an existing installation, change service URLs and reader credentials in Settings. The generated commands configure reader authentication before preflight.
| Setup | What you gain | What runs |
|---|---|---|
| Basic | A complete edit from metadata and small CPU picture classifiers | The app |
| GPU | Picture descriptions, more context for selection and extra sharing checks | The app, inference and a caption service, with Laya ready |
| Full | GPU features plus a text reader's small corrections to the draft | The GPU setup and an explicitly enabled reader |
A video encoder speeds up rendering. It does not enable the GPU selection tier. Hardware encoding and the render worker are separate choices.
Use tier: basic to select it explicitly. GPU and Full values are unchanged.
See Can I run this? for exact platform evidence, candidate versions and untested routes.
What you give up with Basic
| Feature | Basic | GPU | Full |
|---|---|---|---|
| Dates, favourites, people, trips, rules edit, titles and bundled music | Included | Included | Included |
| Picture descriptions for selection | No caption model | Included | Included |
| Additional document/sensitive-content detectors and Laya audience check | Not used | Included | Included |
| Text-reader refinement of the draft | Not used | Not used | Included; reports factual fallback if a reading fails |
| Maximum output resolution | 1080p | 4K when the rendering setup supports it | 4K when the rendering setup supports it |
| Generated music and stem mixing | Optional, if configured and memory permits | Optional | Optional |
Basic keeps the complete editing pipeline. It omits the extra picture interpretation, audience check and reader refinement. Model checks are not a guarantee that every selected picture is suitable for sharing.
Rendering has its own limits. HDR, 60 fps, hardware encoding and animated title effects depend on the encoder, renderer and output settings. A Basic Mac can use its rendering GPU while keeping Basic selection and the 1080p cap. A small CPU-only NAS uses static title plates with fades instead of moving kernel effects. Generated music is not a Full-only feature: Basic can run it locally too.
Measured setups record cold runs across a NAS, a shared Kubernetes GPU and two Macs, including the extra render time and file size for 4K HDR, and name the features exercised. Basic remains 1080p on both NAS and Mac.
Basic: start with the film
The app reads your library's dates, favourites, people and locations, prepares picture facts on its CPU, and builds the edit. Template titles and bundled music are included. There is no model server to run and no hosted AI subscription to buy.
Fast path: follow the Quick start, then the After install steps. Docker Compose needs two CPU cores, 4 GiB free RAM and about 25 GB for persistent data, plus room for the image and finished films. NAS notes cover permissions and access from your desktop.
Cost: the first cut prepares the pictures in its period. Later cuts reuse compatible facts; they still render the video. Basic output stops at 1080p. See measured setups for render time once facts are already prepared; that is not a first-install time.
Check: capabilities should report Basic. A software encoder is a valid result. Run
preflight after models fetch and resolve errors before cutting a month.
GPU: understand more of the pictures
Captions add information beyond dates and faces: what is happening, what objects are present, and how a picture fits the story. The GPU tier also adds document and sensitive-content checks and a family-viewing pre-screen. Review the cut before sharing it; model checks can miss things. See the Basic vs GPU example for one CC0 month's changed cut and its limits; it does not claim that a model always improves a film.
Fast path: on Apple Silicon, use the native Mac setup.
The all-mac extra makes Metal available to automatic tier detection; captions still need a
server. On a NAS, use GPU inference and
captions on a machine you control. The
combined GPU service can host both.
Cost: more model downloads and preparation work, plus memory for the model services. For 4K output, allow at least 8 GB for the app in addition to those services. A second machine adds network transfer and another service to maintain; a GPU does not remove the first preparation.
Check: capabilities should report GPU. preflight must also find the caption service and
Laya runtime/checkpoint. Detecting a GPU alone does not prove either is ready.
Full: refine the draft
The rules editor still builds the film. A text model reads that draft and can propose small changes, such as replacing a weak shot or tightening a story. Each change must pass the selection checks. If the reader cannot answer, the app keeps the rules draft and reports why. Watch the separate Basic and Full example for one observed cut; it is not a quality guarantee.
Fast path: start with the GPU setup, then enable a text reader.
Use its exact served model name and URL, and explicitly set llm.enabled to true.
Entering a URL alone does not turn the reader on. Native Mac and Linux installs can run the
app-owned reader locally; Docker and Kubernetes use an external reader server.
Cost: another model's memory and time, or the hosted provider's fees. The reader receives annotation text, including people and place names. It does not receive the original pictures for this edit pass. Privacy lists each service and what it receives.
Check: capabilities should report Full. preflight checks the reader as well as the GPU
requirements. Compare the same month before deciding whether the changes are worth it to you.
What a model adds explains which edits it may propose.
A reader on NAS
You can enable a reader for written titles and music mood while keeping Basic selection. It does not enable the Full edit pass or sentence films by itself. Generated music is another optional service; bundled music already works.
Pick the platform
| You already have | Install path |
|---|---|
| Linux with Docker | Docker Compose |
| Synology or another NAS | NAS setup |
| Apple Silicon Mac | Native Mac setup |
| Kubernetes cluster | Kubernetes |
Measured setups covers installation checks, picture preparation and finished films separately. It names missing measurements too. Rendering benchmarks alone do not prove that a new user can install and finish a film.
Change services later
Use Settings for inference and caption URLs, and the reader's URL, model and enabled switch. If a field is locked, its source tells you which environment variable or YAML value controls it; move that value into Settings first. Compatible picture facts and your review decisions survive a tier change.
Check after each change:
immich-memories preflight
immich-memories capabilities
For Docker, prefix each command with docker compose exec immich-memories.
Then cut the same month again and compare the results.