Skip to main content

What a GPU or a model adds

Immich Memories makes the whole film on a plain NAS: one container, one models fetch, no GPU, no model to host. A GPU and a text model each add features on top of that. Nothing is lost when you add one later, because everything the app works out about a picture is banked and reused.

The app picks its setup by itself (tier: auto): a plain NAS by default, GPU once it finds a GPU to read pictures on (in this box or through the inference service), Full when a text model is configured as well. The GPU and Full tiers then expect the caption server, and immich-memories preflight says when it is missing.

Feature by feature​

FeaturePlain NAS+ GPU+ GPU and a model (Full)
Choosing the picturesThe rules editor builds the film from dates, places, favourites, the people Immich knows, and a picture encoder with eight small classifiers and two detectors, all on the CPUSame draftThe model reads the finished draft in blocks of 12 shots, names the ones that add nothing and swaps in better pictures of the same moments
A description of each pictureNoneA one-line caption for every picture in the cut and every candidate to replace oneSame, and the model reads them
Family-viewing checkRules and a sensitive-content detector hold back what isn't for sharingA second reader checks the captions and can hold more pictures back (never fewer)Same as GPU. The text model never decides what is shareable
The titleBuilt from the dates, the people and the year, or the album, holiday or trip nameSameWritten by the model from what the film holds
The musicA bundled track, calm by defaultSameThe model picks the mood, tempo and genre from what the cut is about
Reading the picturesOn the CPU, once per picture, then bankedThe encoder and classifiers run on the GPU: the same answers, soonerSame as GPU

A text model on its own, with no GPU, still writes titles and picks the music mood. Choosing the pictures stays on the NAS rules until the GPU tier is there too, because the model's polish reads the captions.

What each one needs​

  • A GPU: an NVIDIA card or a Mac, either in this box or running the inference service on another machine, plus the caption server.
  • A model: any OpenAI-compatible text model with a 32k context, local (such as Gemma 4 E4B) or hosted: Add a reader.

Two more, separate from the tiers​

  • Encoding and title effects: any GPU the render can reach (an Intel or AMD iGPU through VA-API or Quick Sync, NVIDIA through NVENC, a Mac through VideoToolbox) encodes the film faster and draws the animated title effects. Without one, the CPU encodes and the titles keep their text and timing without the effects: Hardware encoding. A render worker moves the whole render to an NVIDIA box.
  • Generated music: an original track for each film instead of a bundled one, from ACE-Step or MusicGen: Generated music.

What each add-on sends where is on Privacy, and the time and memory each one costs is on Measured.