Skip to main content

What a model adds, what it costs

Immich Memories makes the whole film on a plain NAS. The add-ons below make it better or faster, and each one plugs into the same install. Everything they work out is banked next to what the NAS already knows, so switching one off later loses nothing.

What each one adds, feature by feature, is on What a GPU or a model adds. This page is the practical side: what each add-on needs, and what it sends where.

Swipe sideways, or focus the diagram and use the arrow keys, to see more.

The add-ons​

Add-onWhat it buysWhat it needsWhat leaves the box
A readerTitles and music mood on every tier; on Full, an account of the period and refinement of the NAS draftA text model with a 32k context, such as local Gemma 4 E4B. Selection refinement also needs GPU capability, captions and LayaThe candidates' annotation lines, people and place names included, to the model. Never a picture
CaptionsDescriptions for selected pictures and replacement candidates, used by the reader and LayaThe supported 500M vision model, or explicit opt-in to a vision-capable LLMA 400 px tile of each requested picture, once per caption generation, to the chosen provider
Inference on a GPU boxThe encoder, its eight heads and the two detectors on a card or a bigger CPUA second machine, CPU or NVIDIAA preview of each picture, once, to your service
A render workerThe encode on a GPU box instead of the NASAn NVIDIA box running the same app versionThe chosen cut and your Immich key; the worker fetches the originals itself
Generated musicAn original track per film instead of a bundled oneACE-Step on a Mac or an NVIDIA box (7 to 29 GB free for its weights), or a MusicGen serverA text prompt (mood, tempo, length) to your music server

Every destination defaults to localhost or off. Pointing one at another host is the consent step, and Privacy lists every switch.

How the model changes the cut​

On the full tier the rules editor still builds the draft. The model reads it, writes an account of the period, and polishes it: it names the shots that add nothing and swaps in better pictures of the same moments, while favourites, close family and the family-viewing holds stay put. If the model can't answer, the rules draft ships and the log says why. How the polish decides, with diagrams: What a model adds. Which setup reaches full: The three tiers.

What it costs​

Time, memory and euros per setup belong on Measured, with the code revision, hardware and cache state. Plan for these costs:

  • The reader and caption server need memory alongside the app. Reader memory also depends on context length and concurrent requests, not just the model's weight size.
  • NAS does not require a caption server. GPU and Full need a caption provider and reuse existing captions.
  • A hosted reader bills tokens, and the prose is banked, so a week is paid for once, not once per film.