Skip to main content

Choose your setup

The Basic tier makes a complete film: stories, favourites, people, trips, time order, titles and bundled music. It runs on a NAS, a Mac or a cluster without GPU inference. Start there. Add services for a change you want to see in the film.

Setup files for docs build v1.0.0-rc.1. The app image and release assets use this same version.

Create a key with the ten required read permissions. Add the five upload permissions only to send films back; asset.delete is optional. Avoid the All preset.

Choose a tier

Can I run this? Check the version and topology matrix.

Not yet tested as an end-to-end generated installation on Linux Docker Compose. Files are checked with Compose/Kustomize and the form is checked in a browser. Those checks do not run this installation. Earlier NAS, Mac and GPU Kubernetes checks are recorded in the measured results. Tried it? Report your platform, release and preflight result.

Your choices produce files in this browser. Nothing is sent to an Immich or model server.

Enter your Immich API key.

Pick the lane you already run. Every lane gets the same app.

For an existing installation, change service URLs and reader credentials in Settings. The generated commands configure reader authentication before preflight.

SetupWhat you gainWhat runs
BasicA complete edit from metadata and small CPU picture classifiersThe app
GPUPicture descriptions, more context for selection and extra sharing checksThe app, inference and a caption service, with Laya ready
FullGPU features plus a text reader's small corrections to the draftThe GPU setup and an explicitly enabled reader

A video encoder speeds up rendering. It does not enable the GPU selection tier. Hardware encoding and the render worker are separate choices.

Use tier: basic to select it explicitly. GPU and Full values are unchanged.

See Can I run this? for exact platform evidence, candidate versions and untested routes.

What you give up with Basic​

FeatureBasicGPUFull
Dates, favourites, people, trips, rules edit, titles and bundled musicIncludedIncludedIncluded
Picture descriptions for selectionNo caption modelIncludedIncluded
Additional document/sensitive-content detectors and Laya audience checkNot usedIncludedIncluded
Text-reader refinement of the draftNot usedNot usedIncluded; reports factual fallback if a reading fails
Maximum output resolution1080p4K when the rendering setup supports it4K when the rendering setup supports it
Generated music and stem mixingOptional, if configured and memory permitsOptionalOptional

Basic keeps the complete editing pipeline. It omits the extra picture interpretation, audience check and reader refinement. Model checks are not a guarantee that every selected picture is suitable for sharing.

Rendering has its own limits. HDR, 60 fps, hardware encoding and animated title effects depend on the encoder, renderer and output settings. A Basic Mac can use its rendering GPU while keeping Basic selection and the 1080p cap. A small CPU-only NAS uses static title plates with fades instead of moving kernel effects. Generated music is not a Full-only feature: Basic can run it locally too.

Measured setups record cold runs across a NAS, a shared Kubernetes GPU and two Macs, including the extra render time and file size for 4K HDR, and name the features exercised. Basic remains 1080p on both NAS and Mac.

Basic: start with the film​

The app reads your library's dates, favourites, people and locations, prepares picture facts on its CPU, and builds the edit. Template titles and bundled music are included. There is no model server to run and no hosted AI subscription to buy.

Fast path: follow the Quick start, then the After install steps. Docker Compose needs two CPU cores, 4 GiB free RAM and about 25 GB for persistent data, plus room for the image and finished films. NAS notes cover permissions and access from your desktop.

Cost: the first cut prepares the pictures in its period. Later cuts reuse compatible facts; they still render the video. Basic output stops at 1080p. See measured setups for render time once facts are already prepared; that is not a first-install time.

Check: capabilities should report Basic. A software encoder is a valid result. Run preflight after models fetch and resolve errors before cutting a month.

GPU: understand more of the pictures​

Captions add information beyond dates and faces: what is happening, what objects are present, and how a picture fits the story. The GPU tier also adds document and sensitive-content checks and a family-viewing pre-screen. Review the cut before sharing it; model checks can miss things. See the Basic vs GPU example for one CC0 month's changed cut and its limits; it does not claim that a model always improves a film.

Fast path: on Apple Silicon, use the native Mac setup. The all-mac extra makes Metal available to automatic tier detection; captions still need a server. On a NAS, use GPU inference and captions on a machine you control. The combined GPU service can host both.

Cost: more model downloads and preparation work, plus memory for the model services. For 4K output, allow at least 8 GB for the app in addition to those services. A second machine adds network transfer and another service to maintain; a GPU does not remove the first preparation.

Check: capabilities should report GPU. preflight must also find the caption service and Laya runtime/checkpoint. Detecting a GPU alone does not prove either is ready.

Full: refine the draft​

The rules editor still builds the film. A text model reads that draft and can propose small changes, such as replacing a weak shot or tightening a story. Each change must pass the selection checks. If the reader cannot answer, the app keeps the rules draft and reports why. Watch the separate Basic and Full example for one observed cut; it is not a quality guarantee.

Fast path: start with the GPU setup, then enable a text reader. Use its exact served model name and URL, and explicitly set llm.enabled to true. Entering a URL alone does not turn the reader on. Native Mac and Linux installs can run the app-owned reader locally; Docker and Kubernetes use an external reader server.

Cost: another model's memory and time, or the hosted provider's fees. The reader receives annotation text, including people and place names. It does not receive the original pictures for this edit pass. Privacy lists each service and what it receives.

Check: capabilities should report Full. preflight checks the reader as well as the GPU requirements. Compare the same month before deciding whether the changes are worth it to you. What a model adds explains which edits it may propose.

A reader on NAS​

You can enable a reader for written titles and music mood while keeping Basic selection. It does not enable the Full edit pass or sentence films by itself. Generated music is another optional service; bundled music already works.

Pick the platform​

You already haveInstall path
Linux with DockerDocker Compose
Synology or another NASNAS setup
Apple Silicon MacNative Mac setup
Kubernetes clusterKubernetes

Measured setups covers installation checks, picture preparation and finished films separately. It names missing measurements too. Rendering benchmarks alone do not prove that a new user can install and finish a film.

Change services later​

Use Settings for inference and caption URLs, and the reader's URL, model and enabled switch. If a field is locked, its source tells you which environment variable or YAML value controls it; move that value into Settings first. Compatible picture facts and your review decisions survive a tier change.

Check after each change:

immich-memories preflight
immich-memories capabilities

For Docker, prefix each command with docker compose exec immich-memories. Then cut the same month again and compare the results.