Skip to main content

Pixel evidence and preparation

This page lists what preparation reads, which pictures it covers, and when each producer runs.

Scope and timing​

A film acquires cheap facts for its reach: the pictures it could select (for a person film, exactly the pictures Immich recognised that person on, strict per picture), the other stills of their Live Photo bursts, and every picture of the same five-minute capture run, because the exposure rule reads the whole run. The rest of the window is read as Immich metadata only, since moments and episodes are cut from all of it. A cut that selects a picture it never prepared stops rather than ship it. Captions and clip checks wait until after the rules draft, for selected shots and actual candidates. A reader may use the selected shot's whole episode for context without captioning every neighbour. immich-memories prepare reads a whole scope ahead of time when explicitly requested.

In order: resolve the reach, acquire the required pixel facts, build the rules draft, caption and inspect the selected candidates, then bank the complete evidence.

Admission refuses a few things before anything is read: a video over five minutes (advanced.analysis.max_source_video_seconds, 300 s), the video half of a Live Photo (it plays inside its still), anything tagged immich-memories/generated or listed in this install's upload receipts (a film this app made is not footage), and pictures that look forwarded rather than shot on your camera. After the heads run, screenshots and photos of screens go too: a phone-screen pixel size, the screen head, or the document detector, enabled on GPU and Full, calling it a screenshot, a table or a QR code.

The eight heads are small classifiers over one pinned DINOv2 encoder: location, people, children, activity, venue, frame_kind, screen and uncovered_person. The two detectors are nsfw_marqo (exposure) and doc_docling (documents), enabled on GPU and Full only. Basic keeps the eight heads, including screen, frame-kind and uncovered-person checks. Every fact is banked in the store under its producer's version, so the next cut asks nothing twice.

Preparation follows the resolved product tier (tier: auto by default):

TierWhat reads the pixelsWhen you get it
basicpreviews, pixel facts, face boxes, the eight headsno usable local GPU or GPU inference service
gpuBasic facts, Marqo and Docling, plus missing captions and clip evidence for selected shots and candidates; Laya reads their captionsGPU inference without a configured prose LLM
fullthe same pixel producers as GPU; a prose LLM reads annotation text to refine selectionGPU inference and a configured prose LLM

Basic needs immich-memories models fetch once. A configured LLM alone does not change selection from Basic, but can still write titles and music mood. GPU and Full enable captions by default; Basic can use a vision-capable LLM only with explicit caption-provider opt-in. Captions are an add-on: Add captions.

All selection internals.