Text reader and provider contracts
For the working local and hosted recipes, start with Add a text reader. This reference covers the audience classifier and provider transport. The selection reader receives text; image captions require a separate explicit choice.
Reader operating modes
advanced.llm.enabled defaults to false; a model, URL or key alone never enables it. With it enabled, an empty base_url with openai-compatible or ollama selects the app-owned llama.cpp process; a nonempty URL selects an external server. The default local model is gemma-4-E4B-it-Q4_0. For the install recipe, use Add a reader.
advanced:
llm:
enabled: true
base_url: ""
model: gemma-4-E4B-it-Q4_0
local_server: llama-server
local_context: 32768
Owned inference supports Linux and macOS. models fetch downloads the pinned default GGUF and projector when this local reader is configured. A custom model names a GGUF path; local_mmproj supplies its projector. Preflight checks the executable and files without loading the model.
The app loads the owned reader on demand and waits for active requests before stopping it to release memory for local ACE-Step or Demucs. At the render boundary, selection also closes its Laya scorer, stops the owned reader and clears unused local-runtime buffers. The next reader call loads it again. Cancellation waits for native audio work to finish or time out before releasing its memory lease. The owned reader always serves one request at a time (reader_concurrency cannot increase it). Docker and Kubernetes need an external server; the app image has no llama-server. An external server owns its own model lifetime: the app cannot assume it is safe to unload it for other clients.
immich-memories capabilities
immich-memories capabilities --verify-local
--verify-local uses synthetic inputs and installed weights to check the configured owned reader and local audio path. It does not certify external services or a complete film, and cannot be combined with --test-music, which may download models.
Local server defaults
For an external openai-compatible server, the app classifies the base_url host. Loopback,
private IP addresses and single-label service names count as local: one request at a time, with
prompt-only episode JSON and a repetition penalty of 1.0. A dotted hostname counts as hosted:
four concurrent requests, episode schemas enabled and no repetition penalty by default. For a
LAN server with a dotted name, set reader_concurrency: 1 and structured_output: false if its
constrained decoding stalls; the latter disables schemas for every request, not just episode
readings.
The Laya audience pre-screen
Laya answers the sharing question locally: does the caption describe
a bath, a nappy change, breastfeeding or one of the other
private activities a family film holds back. Laya is a 0.4B text classifier (Apache-2.0),
fine-tuned on captions of public CC BY photographs whose authors are credited in the archive. It
reads the ingest caption, in about 14 ms a shot on Apple silicon. It works with the rules reader
and the prose reader when preparation produces captions. The gpu and full tiers enable it;
the default basic tier uses the picture classifiers and rules.
uv tool install --python 3.12 "immich-memories[all-mac]" --with laya-mlx # Apple Silicon uv-tool install
immich-memories models fetch --laya # platform-specific, digest-pinned
laya-mlx is not included in an app extra. In a checkout, use
uv pip install --python .venv/bin/python laya-mlx; with pip, run
python -m pip install laya-mlx using the app's Python. A bare system pip install does not
add it to a uv tool environment. Without the runtime, sharing reports that heads and rules
are being used alone.
models fetch chooses the Apple archive on Apple silicon and the portable ONNX archive on
Linux, Windows and Intel Macs. ONNX needs the editorial extra for CPU or editorial-cuda for
NVIDIA. The Apple download is 811 MiB (851 MB). The ONNX download is 836 MiB (877 MB)
and expands to about 1.70 GB;
its calibrated default threshold is
0.185, while MLX keeps 0.186. Both archives are SHA-256 checked. A download mirror must keep the
archive's filename. When configuring a checkpoint for a different backend manually, set its
threshold explicitly too.
The GPU and Full product tiers enable Laya automatically. A saved laya_audience value cannot
override the resolved tier. Fetch the checkpoint and check the services with immich-memories preflight.
It only adds holds. The detector holds (the sensitive-content detector and the uncovered-person
head) apply first and are never lifted, its findings go through the same support checks as the
reader's, and a shot it doesn't answer stays held to the family. Sharing never asks the prose
LLM, including when Laya is absent or a detector flags exposure. The MLX threshold,
laya_audience_threshold: 0.186, kept every hold of its public calibration split. Its known gap:
a travel or administrative document (a boarding pass, an invoice) can slip
through, since few such captions were in its training data. The detectors stay the floor either way.
To use an extracted ONNX checkpoint, point advanced.editorial.laya_checkpoint at the
directory containing model.onnx, model.onnx.data, rl_agent_config.json and tokenizer/.
The scorer chooses CUDA when available and otherwise uses CPU. This path needs neither
PyTorch nor MLX.
For the audience ONNX export, set laya_audience_threshold: 0.185. This threshold was
chosen on the public calibration split to retain all 15 MLX holds. On 3,143 held-out
captions it retained all 23 MLX holds and added one. These are classifier checks, not
end-to-end film performance; the calibration numbers stand on their own.
Ollama
Native Ollama uses /api/generate. For Gemma 4 E4B, set the context in options
and explicitly disable thinking for the text reader:
advanced:
llm:
enabled: true
provider: ollama
base_url: http://localhost:11434
model: gemma4:e4b-it-q4_K_M
extra_params:
think: false
options:
num_ctx: 32768
Pull that model with ollama pull gemma4:e4b-it-q4_K_M. The explicit tag records
the quantization tested. extra_params.think: false overrides thinking on every
native request, including titles; omit it if you deliberately want reasoning and
allow enough output tokens for it. Without it, thinking can consume part of the
first large episode request's token budget, leaving the JSON answer truncated.
For Ollama's OpenAI route, use its own reasoning switch:
advanced:
llm:
enabled: true
provider: openai-compatible
base_url: http://localhost:11434/v1
model: gemma4:e4b-it-q4_K_M
no_thinking_params:
reasoning_effort: none
That route cannot set the context per request. Start the server with
OLLAMA_CONTEXT_LENGTH=32768 ollama serve, or create a model with
PARAMETER num_ctx 32768 in its Modelfile. The default
chat_template_kwargs.enable_thinking: false does not disable thinking on
Ollama's compatible route. See Ollama context length
and OpenAI compatibility.
From a container, use the host's reachable address instead of localhost.
Both recipes passed the larger episode and story-selection probes on an M5 Max. The native recipe passed 33/34 synthetic feature checks; the compatible recipe passed 32/34. Motion failed on both; the compatible route also missed a period-summary fact. Measure your setup has the numbers. These settings were tested with Gemma 4 E4B, not every model Ollama can serve.
Providers and dialects
Five provider values, three code paths. ollama speaks Ollama's native API, anthropic speaks
/v1/messages, and openai-compatible and openai speak /v1/chat/completions, so anything
serving that endpoint works: mlx-vlm, oMLX, vLLM, Ollama's compatibility layer, Groq, OpenAI
itself. zai is the anthropic adapter with z.ai's URL and reasoning level filled in, and it is
the one provider that picks its adapter from the base_url path, because z.ai serves both dialects
on one host: .../api/anthropic gets /v1/messages, .../api/paas/v4 gets /chat/completions.
openai, anthropic and zai fill a blank base_url with https://api.openai.com/v1,
https://api.anthropic.com and https://api.z.ai/api/anthropic, respectively. Set an explicit URL for other external servers. Blank openai-compatible or ollama uses an owned local reader. The named providers supply their reasoning dialect where settings retain their defaults; explicit request parameters take precedence.
The Messages API path is POST {base_url}/v1/messages with x-api-key,
anthropic-version: 2023-06-01 and the prompt as one user message. Nothing about it is
Claude-specific: point base_url at whoever serves the dialect. Answers come back as a JSON
envelope the app validates itself, and no provider-side JSON mode is used, so a host without one
loses nothing.
advanced:
llm:
enabled: true
provider: "anthropic"
model: "claude-haiku-4-5"
api_key: "${ANTHROPIC_API_KEY}"
thinking: "auto"
Start with thinking: "auto": it omits both reasoning switches and uses the model's default. Haiku 4.5 does not accept adaptive thinking. Sonnet 5.5, Opus 5.5 and Fable models reject thinking: {type: "disabled"}. The app's other thinking levels use adaptive thinking for titles and send the disabled switch on bulk calls; use those only with a model that accepts both. See Anthropic's thinking matrix.
The preset drops temperature. An explicit budget_tokens override alone does not remove the preset's adaptive effort field. Keep thinking: "auto" unless you have checked every custom field against your model's contract.
advanced:
llm:
enabled: true
provider: "zai"
base_url: https://api.z.ai/api/anthropic
model: "glm-5.3-flash"
api_key: "${ZAI_API_KEY}"
thinking: "low"
Set base_url: https://api.z.ai/api/anthropic for its Messages API; use the endpoint your account supports. The
other route, https://api.z.ai/api/paas/v4, is the OpenAI-compatible one and answers that account
429 code 1113, Insufficient balance; set it explicitly if your account is the other kind. The
GLM-5 line refuses disabled, so the preset sends low.
For any other host serving the Messages API, set base_url yourself and use thinking: "auto",
which sends no reasoning field and takes the host's default.
Reasoning
On a server whose chat template reasons by default, a bulk call reasons through its small token
budget, stops mid-thought and returns nothing parseable. llm.no_thinking_params stops that, and
its default is already the Qwen dialect:
llm:
thinking: "disabled" # default
no_thinking_params: # merged into every non-thinking call
chat_template_kwargs:
enable_thinking: false
A server that reasons only when asked wants no_thinking_params: {} instead. low, high or
max switch title generation to reasoning while other calls use the non-thinking policy.
thinking_params carries the fields that title call sends; OpenAI's reasoning models want {"reasoning_effort": "medium"} there, which provider: openai fills in. The provider's own switch
is merged in even when you set your own params, and a thinking key you write yourself wins.
Ollama has neither chat dialect: its switch is a bare top-level think, billed inside
num_predict. A load-bearing call gets think: true, a bulk call gets no switch at all (a model
without a thinking mode answers think with a 400), and a server that reasons unasked is learned
from its first thinking block: every later call then gets 16,384 extra tokens in num_predict. An
extra_params.options.num_predict you set yourself wins.
A level is a request, not a promise. z.ai's .../api/anthropic route answers HTTP 200 to every
setting and then reasons on its own terms, so on that route the reader reads the first text block
and skips the reasoning in front of it, asks for 16,384 tokens on top of the caller's cap, and turns
a reply with no text block into an error naming the stop_reason. A provider's own error code
and message go into the log line, cut at 300 characters.
Structured replies
The default selects the mode by request type. Free-text questions, titles and period accounts
request their JSON schemas on local and hosted endpoints. External local episode readings use prompt-only
JSON because oMLX can stall on their nested schema. The app-owned llama.cpp reader and hosted
episode readings retain the schema.
Both modes work against the same endpoint in one process. advanced.llm.structured_output can
explicitly enable or disable structured output for that endpoint.
Occasion discovery, holiday checks and special-day decisions request JSON object mode.
Native Ollama receives format: "json" for these calls. Plain-text requests do not inherit
that mode from an earlier decision.
If a provider refuses schema mode and asks for json_object, the app retries once in object
mode and carries the schema in the prompt. It remembers that choice for the endpoint and model
for the rest of the process and logs the adaptation once. A refusal of the whole response-format
parameter removes that parameter. An invalid schema or another ordinary HTTP 400 still fails.
For a provider already known to lack schema support, structured_output: false skips negotiation.
Batch mode
Episode readings are independent prompts. Supported external providers can submit them as batches; pricing and completion latency depend on the host. Owned local inference does not use batch APIs.
advanced:
llm:
batch: "auto" # off (default) | auto
batch_min_requests: 8 # below this, asking one at a time is quicker
batch_max_wait_minutes: 60 # then ask whatever is left in real time
It pays on an unattended run (the nightly auto run, a prepare --overviews over a year) and not
on a run someone is waiting for: a batch is queued work, and "usually within minutes" is not a
promise you want between a click and a film. If it goes wrong you lose the discount and nothing
else. Anything unanswered by batch_max_wait_minutes, any line the provider refused and any
answer the parser won't read is asked again in real time.
| Provider dialect | Batch route |
|---|---|
| OpenAI-compatible | /v1/batches |
| Anthropic-compatible | /v1/messages/batches |
| Ollama native | No batch route |
The endpoint must support the route. A compatibility dialect alone does not establish that it does, and pricing comes from the provider. Owned local inference stays in real time.
The route is probed once before anything is queued. A host that doesn't serve it is asked once, logs why, and reads in real time for the rest of the run.
What the software sends
- Prompts are bounded before they go out: episode reads at 24,000 characters and 90 pictures a page, story synthesis at 32,000, the period account split into pages. The owned reader defaults to a 32k context; context memory also depends on the server and concurrent work.
- Answers are parsed against the stage's contract. An answer the contract refuses costs one
repair round on that call. A moment pick still refused after its repair doesn't end the film:
those rows get the moments the no-model film would pick, and the story's pick record says why
(
pick-rules-fallback). - A reply cut off at its token cap keeps what it finished. The episode readings or period accounts it wrote whole are kept, and only the unfinished ones are asked again. An episode asked again gets the full 4,000-token ceiling rather than its own estimate.
- The selection reader receives text only. Image captions use the same settings only with an explicit LLM-caption opt-in.
For endpoint validation, use the provider conformance checks below. Record usage and timings separately from the quality of a finished film; measurement guidance explains the comparison.
Provider conformance
From a checkout, run:
make llm-conformance CONFIG=/path/to/provider.yaml OUTPUT=/tmp/llm-conformance
The command reads only advanced.llm (or llm) from that file. It sends synthetic evidence
through production features and can incur provider charges. Free-text cases also read the
public WordNet corpus installed by immich-memories models fetch. Each banked feature gets a
fresh temporary SQLite store. The suite opens no personal library or people file.
The 34 probes cover occasion, trip and people titles; free-text reading, linking and pool selection; occasion discovery; music mood; image captions and motion; full and lean episode readings; month and year accounts; editorial grouping, weighting, picking and recurring activities; and both audience text readers still present in the code.
Each row reports whether the provider was called, HTTP attempts, reported tokens, elapsed seconds, validity and a feature-specific quality check. Unknown usage is shown as unknown. A local fallback fails the check. A failed feature leaves its row and the other checks continue. The command exits with status 1 if a feature fails or a production model call site has no case. The guard test scans the code for model calls and shared prompt adapters; runtime observation also checks that each case reached its declared call sites.
OUTPUT is optional. When set, it receives an incremental JSON report, a Markdown table,
and private request/reply evidence without HTTP headers. Keep these files private: provider
errors can include account details. Use the input, cached-input and completion token counters
with the provider's rates to calculate cost. A run with missing usage gives only a cost floor.
Use the report to compare passes, failures, call counts and reported-token cost for your endpoint. Benchmarking your own setup explains what to record. Measure your setup records tested endpoints, elapsed time, token costs and quality limits, including the motion-direction weak spot.
These checks measure the configured endpoint on small fixtures. They do not replace checking the quality of a complete film or testing asynchronous batch delivery.