Skip to main content

Add a reader

The NAS makes the film without one. A reader is a text model that writes the prose (what happened in each episode, an account of the period, the film's title, the music's mood). On Full it can refine the rules draft: a proposed replacement must pass the shared checks before taking a shot's seat. It can also keep the existing cut. It starts from the NAS draft and never sees a picture. The captioner reads selected shots and replacement candidates; existing captions are reused. The reader can read a selected shot's whole episode for context, without asking the captioner to fill every neighbour first. What exactly it changes, with diagrams: What a model adds.

The same endpoint can separately act as the caption provider only when you explicitly enable LLM captions. That role sends images, needs vision support, and can cost much more, especially hosted. The selection reader still receives text.

What you need​

Which servers and hosted models have been checked, and when: Supported and tested.

  • A text model with at least a 32k context. No vision needed.
  • An endpoint that speaks the OpenAI /v1/chat/completions or the Anthropic /v1/messages API, or Ollama's own.
  • For a local model, a machine that holds it for as long as its server is up. The default is Gemma 4 E4B (mlx-community/gemma-4-e4b-it-6bit on a Mac, google/gemma-4-E4B-it under vLLM or Ollama; Apache 2.0). Allow room for the context cache and caption server as well as the weights. Context length and concurrent requests affect memory use; see Measured.

Gemma 4 E4B is the default. Another model needs the context window and valid JSON responses, but that alone does not establish the quality of its choices. Compare the finished pictures against NAS on the same inputs and settings. A larger model is not an automatic upgrade. Current measurements belong on Measured; older whole-period-reader timings do not describe the bounded refinement path.

With tier: auto, a configured reader and GPU inference select Full. The caption service and Laya must also be ready. Without GPU inference, selection stays on NAS and the app explains what is missing; the LLM can still supply titles and music mood. It is never used automatically as a captioner. See Requirements and tiers.

Local, on a Mac​

oMLX serves MLX models over an OpenAI-compatible API (macOS 15+):

brew tap jundot/omlx https://github.com/jundot/omlx
brew install jundot/omlx/omlx
omlx start # serves on port 8000

Pull mlx-community/gemma-4-e4b-it-6bit from http://localhost:8000/admin/chat, then point the app at it:

advanced:
llm:
provider: openai-compatible
base_url: http://localhost:8000/v1
model: gemma-4-e4b-it-6bit

model must be exactly what the server reports at GET /v1/models. base_url defaults to http://localhost:8080/v1, the app's own port, so always set it. From the app in Docker the host is host.docker.internal, and from a NAS it is the Mac's LAN name: Reaching a model server. A server that answers 401 wants its token in llm.api_key (IMMICH_MEMORIES_LLM__API_KEY).

On Linux with a card, serve the same model with vLLM or Ollama and set base_url and model the same way. mlx-vlm is another way to serve it on a Mac.

Hosted​

Same contract, a provider URL and a key. Read this first: the candidates' annotation lines leave your network, with the people and place names on them, the dates and the captions. No picture does. If that text should stay home, run the model locally.

advanced:
llm:
provider: "openai" # ollama | openai-compatible | openai | zai | anthropic
model: "gpt-4.1-mini"
api_key: "${OPENAI_API_KEY}"

Leave base_url unset and openai, anthropic and zai fill in their own. The prose is banked, so a week is read and paid for once, not once per film. Every outbound request is on Privacy.

Check it​

immich-memories preflight

The LLM row checks that the endpoint answers for your model (Ollama's tag list, a minimal chat call on an OpenAI-compatible host, the model list or a one-token ask on an Anthropic one). It checks a configured LLM on NAS and GPU too, because titles and music mood can use it even when selection uses rules. With no model configured it reads SKIPPED. A model reader with a blank model stops with editorial runtime needs a nonblank LLM model.

A reader that fails mid-film does not fail the film. The period account is asked twice; after the second failure the rules draft ships with the passes a no-model film gets, and the log says so: The model polish did not run (<reason>); the film is the rules draft.

The Laya audience pre-screen​

Laya answers the sharing question locally: does the caption describe a bath, a nappy change, breastfeeding or one of the other private activities a family film holds back. Laya is a 0.4B text classifier (Apache-2.0), fine-tuned on captions of public CC BY photographs whose authors are credited in the archive. It reads the ingest caption, in about 14 ms a shot on Apple silicon. It works with the rules reader and the prose reader when preparation produces captions. The gpu and full tiers enable it; the default nas tier uses the picture classifiers and rules.

pip install laya-mlx                         # Apple Silicon only
immich-memories models fetch --laya # platform-specific, digest-pinned

models fetch chooses the Apple archive on Apple silicon and the portable ONNX archive on Linux, Windows and Intel Macs. ONNX needs the editorial extra for CPU or editorial-cuda for NVIDIA. The Apple download is 811 MB. The ONNX download is 877 MB and expands to 1.70 GB; its calibrated default threshold is 0.185, while MLX keeps 0.186. Both archives are SHA-256 checked. A download mirror must keep the archive's filename. When configuring a checkpoint for a different backend manually, set its threshold explicitly too.

The GPU and Full product tiers enable Laya automatically. A legacy laya_audience setting cannot override the resolved tier. Fetch the checkpoint and check the services with immich-memories preflight.

It only adds holds. The detector holds (the sensitive-content detector and the uncovered-person head) apply first and are never lifted, its findings go through the same support checks as the reader's, and a shot it doesn't answer stays held to the family. Sharing never asks the prose LLM, including when Laya is absent or a detector flags exposure. The MLX threshold, laya_audience_threshold: 0.186, kept every hold of its public calibration split. Its known gap: a travel or administrative document (a boarding pass, an invoice) can slip through, since few such captions were in its training data. The detectors stay the floor either way.

To use an extracted ONNX checkpoint, point advanced.editorial.laya_checkpoint at the directory containing model.onnx, model.onnx.data, rl_agent_config.json and tokenizer/. The scorer chooses CUDA when available and otherwise uses CPU. This path needs neither PyTorch nor MLX.

For the audience ONNX export, set laya_audience_threshold: 0.185. This threshold was chosen on the public calibration split to retain all 15 MLX holds. On 3,143 held-out captions it retained all 23 MLX holds and added one. These are classifier checks; validation of the complete NVIDIA image is tracked in #1385.

Providers and dialects​

Five provider values, three code paths. ollama speaks Ollama's native API, anthropic speaks /v1/messages, and openai-compatible and openai speak /v1/chat/completions, so anything serving that endpoint works: mlx-vlm, oMLX, vLLM, Ollama's compatibility layer, Groq, OpenAI itself. zai is the anthropic adapter with z.ai's URL and reasoning level filled in, and it is the one provider that picks its adapter from the base_url path, because z.ai serves both dialects on one host: .../api/anthropic gets /v1/messages, .../api/paas/v4 gets /chat/completions.

openai, anthropic and zai fill in the vendor's base URL and reasoning dialect where you left the field at its default. openai-compatible fills in nothing. An explicit base_url always wins.

The Messages API path is POST {base_url}/v1/messages with x-api-key, anthropic-version: 2023-06-01 and the prompt as one user message. Nothing about it is Claude-specific: point base_url at whoever serves the dialect. Answers come back as a JSON envelope the app validates itself, and no provider-side JSON mode is used, so a host without one loses nothing.

advanced:
llm:
provider: "anthropic"
model: "claude-sonnet-5" # or claude-haiku-4-5 for the cheap seat
api_key: "${ANTHROPIC_API_KEY}"
thinking: "high" # disabled | low | high | max | auto

That preset handles two things Claude answers HTTP 400 to otherwise: no temperature goes out (from the 4.7 line on, Claude refuses any sampling parameter), and reasoning is asked for as thinking: {"type": "adaptive"} with the level as output_config.effort. Bulk calls send thinking: {"type": "disabled"}, because a bulk call at a 140-token cap that reasons comes back with no answer in it. A model older than that dialect needs the switch written out: thinking_params: {thinking: {type: "enabled", budget_tokens: 2048}}.

advanced:
llm:
provider: "zai"
model: "glm-5.3-flash"
api_key: "${ZAI_API_KEY}"
thinking: "low"

base_url defaults to https://api.z.ai/api/anthropic, where a coding-plan account is served. The other route, https://api.z.ai/api/paas/v4, is the OpenAI-compatible one and answers that account 429 code 1113, Insufficient balance; set it explicitly if your account is the other kind. The GLM-5 line refuses disabled, so the preset sends low.

For any other host serving the Messages API, set base_url yourself and use thinking: "auto", which sends no reasoning field and takes the host's default.

Reasoning​

On a server whose chat template reasons by default, a bulk call reasons through its small token budget, stops mid-thought and returns nothing parseable. llm.no_thinking_params stops that, and its default is already the Qwen dialect:

llm:
thinking: "disabled" # default
no_thinking_params: # merged into every non-thinking call
chat_template_kwargs:
enable_thinking: false

A server that reasons only when asked wants no_thinking_params: {} instead. low, high or max switch two calls to reasoning (title generation and the special-day question in discover-days) while everything else stays fast. thinking_params carries the fields those calls send; OpenAI's reasoning models want {"reasoning_effort": "medium"} there, which provider: openai fills in. true and false still parse, as high and disabled. The provider's own switch is merged in even when you set your own params, and a thinking key you write yourself wins.

Ollama has neither chat dialect: its switch is a bare top-level think, billed inside num_predict. A load-bearing call gets think: true, a bulk call gets no switch at all (a model without a thinking mode answers think with a 400), and a server that reasons unasked is learned from its first thinking block: every later call then gets 16,384 extra tokens in num_predict. An extra_params.options.num_predict you set yourself wins.

A level is a request, not a promise. z.ai's .../api/anthropic route answers HTTP 200 to every setting and then reasons on its own terms, so on that route the reader reads the first text block and skips the reasoning in front of it, asks for 1,024 tokens on top of the caller's cap, and turns a reply with no text block into an error naming the stop_reason. A provider's own error code and message go into the log line, cut at 300 characters.

Structured replies​

Hosted readers request a JSON schema by default. Local endpoints default to prompt-only JSON because some local grammar decoders stall on these schemas. advanced.llm.structured_output can explicitly enable or disable that request shape.

If a provider refuses schema mode and asks for json_object, the app retries once in object mode and carries the schema in the prompt. It remembers that choice for the endpoint and model for the rest of the process and logs the adaptation once. A refusal of the whole response-format parameter removes that parameter. An invalid schema or another ordinary HTTP 400 still fails. For a provider already known to lack schema support, structured_output: false skips negotiation.

Batch mode​

The episode readings are one prompt per episode, and those prompts don't read each other. Every hosted provider sells that shape cheaper: hand the pile over at once, get it back within the day, pay half.

advanced:
llm:
batch: "auto" # off (default) | auto
batch_min_requests: 8 # below this, asking one at a time is quicker
batch_max_wait_minutes: 60 # then ask whatever is left in real time

It pays on an unattended run (the nightly auto run, a prepare --overviews over a year) and not on a run someone is waiting for: a batch is queued work, and "usually within minutes" is not a promise you want between a click and a film. If it goes wrong you lose the discount and nothing else. Anything unanswered by batch_max_wait_minutes, any line the provider refused and any answer the parser won't read is asked again in real time.

ProviderRouteDiscount
OpenAI/v1/batches (Batch API)50 %, documented
Anthropic, and hosts serving its API/v1/messages/batches (Message Batches)50 %, documented
Melious/v1/batches, same shape as OpenAInone: their docs say batches run at the same per-token rate
z.aianswers 404 on /v1/messages/batchesno batch route; stays real time, with the reason in the log

The route is probed once before anything is queued. A host that doesn't serve it is asked once, logs why, and reads in real time for the rest of the run.

What the software sends​

  • Prompts are bounded before they go out: episode reads at 24,000 characters and 90 pictures a page, story synthesis at 32,000, the period account split into pages. A 32k context holds every one.
  • Answers are parsed against the stage's contract. An answer the contract refuses costs one repair round on that call. A moment pick still refused after its repair doesn't end the film: those rows get the moments the no-model film would pick, and the story's pick record says why (pick-rules-fallback).
  • A reply cut off at its token cap keeps what it finished. The episode readings or period accounts it wrote whole are kept, and only the unfinished ones are asked again. An episode asked again gets the full 4,000-token ceiling rather than its own estimate.
  • Text only. No request to the reader carries a picture; a test fails the build if one does.

The prompt shapes the setup matrix probes readers with are in scripts/reader_probe_prompts/. Time, tokens and euros per reader go on Measured.