Config Reference
Every key with its built-in default. Set one in ~/.immich-memories/config.yaml, as an environment
variable, or from the web UI's settings page (saved to the database). Environment beats the file,
the file beats the database, the database beats these defaults; immich-memories config show says
which one set each key (where a setting comes from).
Tier 2 sections (analysis, hardware, llm, musicgen, ace_step, server, auth,
automation, notifications, triage, editorial, inference, free_text, speech) go under an
advanced: key in the file:
advanced:
analysis:
max_album_assets: 5000
hardware:
encoder_preset: "quality"
Both placements are read and merge key by key: a top-level value wins for the same key, while
other keys under advanced: remain set. Everything else
stays top level. Unknown keys inside a section are ignored; unknown top-level keys and invalid
values fail validation at startup.
Tier
One resolved tier controls both preparation and selection. Leave it automatic for normal use.
tier: auto # auto | basic | gpu | full
tier | Models | Reader | Needs |
|---|---|---|---|
basic | eight shared-DINO CPU heads, no captions | rules | models fetch |
gpu | heads, Marqo, Docling, captions and Laya | rules | a caption server, models fetch |
full | everything in gpu | an LLM polishes the rules draft and writes the prose | enable advanced.llm; owned local model by default, or an explicit base_url and server model |
auto is the default. A healthy inference service reporting CUDA, or a local CUDA or MLX/Metal
runtime, selects gpu. An explicitly enabled LLM alongside that capability selects full. Without GPU
inference capability, selection stays on basic and reports what is missing. A renderer's GPU
does not establish inference capability. The runtime check loads no model weights and sends no
pictures; preflight and acquisition still check the actual producers.
editorial.reader, editorial.preparation.tier and editorial.laya_audience are derived from
the product tier. These values are derived rather than separate user controls. Save omits these
derived settings and keeps automatic resolution automatic when the file moves to another host.
An explicit basic, gpu or full pins a tier for a controlled comparison; it does not install or
start its services. full requires an enabled reader: the app-owned local model or a configured API.
Configured LLM titles and music mood work on every tier. Basic and GPU still select with rules, and sharing never asks the prose LLM. Captions use their own configured service; a text LLM is not an automatic caption fallback. A missing Laya checkpoint or runtime is reported and uses the conservative rules fallback; that is a degraded run, not a verified GPU/full comparison.
Env: IMMICH_MEMORIES_TIER=auto. From a source checkout,
uv run python scripts/tier_settings.py prints what each tier runs with. That helper script is
not part of a pip/uv install; immich-memories capabilities reports the resolved setup there.
Preset
One top-level switch that fills several knobs at once. fast is the CPU-only / NAS profile.
Anything you set yourself wins over the preset.
preset: null # null | fast
fast sets, unless you set them yourself: output.resolution: 1080p, output.codec: h264,
output.quality: fast, hardware.encoder_preset: fast and
title_screens.animated_background: false (static title backgrounds and three-view maps). Music generation is already
off by default and stays wherever you put it.
Env: IMMICH_MEMORIES_PRESET=fast. One-off on the CLI: immich-memories --preset fast generate …
(root option, before the subcommand).
Settings saved in the UI or CLI go to the database. Environment variables and config.yaml still win; the app refuses to save a database value that they would override. Bootstrap keys never go to the database: database.* (read before the store opens), and auth.* and server.* (they decide who can reach the app). Set them in the environment or config.yaml and restart. A saved value cannot contain a ${VAR} reference.
Immich connection
Immich Memories supports Immich v2 and v3. Automatic runtime detection is the default:
immich:
url: "https://photos.example.com"
api_key: "${IMMICH_API_KEY}"
api_version: auto # auto | v2 | v3
Keep api_version on auto for normal use. The client detects and caches the server major for
each runtime client; you do not choose it for each generation. Explicit v2 or v3 is a manual
troubleshooting escape hatch for a proxy or unusual deployment that prevents correct detection.
An override forces that API contract.
Run the read-only immich-memories config test to check credentials and see the resolved API
contract without generating or uploading a memory.
Extra accounts
People can upload to separate Immich accounts. Put additional accounts under accounts, by name. The top-level url and api_key stay the primary account,
and the primary is the only account a film is ever uploaded to.
immich:
url: "https://photos.example.com"
api_key: "${IMMICH_API_KEY}"
native_sharing: false # opt into verified native person identities (3.2+; 3.3 experimental)
accounts: {} # name -> url, api_key, api_version (default: none)
# accounts:
# partner:
# url: "https://photos.example.com"
# api_key: "${PARTNER_IMMICH_API_KEY}"
# api_version: auto
A name is lowercase letters and digits joined by single underscores (partner, grandma_2), and
primary is taken: it means the top-level account. The name is what a person alias bound to that
account records. Configuring an account adds nothing to your films: a run reads only the accounts it
selects, and every selected account has to answer /users/me with its key before anything is read.
generate --accounts primary,partner reads the named accounts into one film on the CLI, and
automation.accounts does the same for the daily scan. config test and
preflight check each one, one line per account, with no key printed:
A second Immich account.
The key is a secret like the primary one: redacted from logs and issue reports, and sealed with
IMMICH_MEMORIES_SECRET_KEY when saved to the database (the whole immich.accounts map is one
encrypted row). From the environment, name the account in the variable:
IMMICH_MEMORIES_IMMICH__ACCOUNTS__PARTNER__API_KEY.
Render worker
The CLI and web UI can send an already selected film to a trusted render worker.
Blank worker_base_url renders on the app's machine.
render:
worker_base_url: ""
worker_token: "" # Or ${RENDER_WORKER_TOKEN}
allow_insecure_http: false # Explicitly accept a non-loopback cleartext HTTP worker
timeout_seconds: 3600 # Wait for rendering and download; maximum 86400; greater than 0, at most 86400 seconds
fallback_to_local: false # Set true to allow local rendering after a worker failure
Use the same app version on both machines. The worker receives the selected assets,
exact cuts, Live source material, titles, locations, audio markers and the Immich API key
so it can download the sources directly. Configure a worker you trust, reachable over
your private network or HTTPS. The handoff request carries that Immich key, so a
non-loopback http:// worker URL is refused until allow_insecure_http: true says you
meant it; loopback addresses and HTTPS need no opt-in. Preflight names the transport
before the first render request. Requests require the worker token and do not follow redirects.
The worker accepts H.264 or H.265 MP4. When a worker is configured, MOV and ProRes
render locally only with fallback_to_local: true; with false, the run fails.
Orientation only sets the canvas; it does not change the selection. Speech detection and cut
selection run before handoff. Music and Immich upload finish on the app after it checks the returned film.
See worker deployment for Docker Compose and Kubernetes examples.
Video analysis
advanced:
analysis:
# Media the camera roll did not shoot (see Configuration → Footage the
# camera roll did not shoot). Setting the list replaces it; [] turns it off.
exclude_filename_patterns: # case-insensitive globs on the source filename
- "RingVideo_*"
- "RPReplay_Final*"
- "Screen Recording *"
- "Screenshot*"
- "img-*-wa[0-9][0-9][0-9][0-9]*"
- "vid-*-wa[0-9][0-9][0-9][0-9]*"
exclude_stills_without_camera_exif: true # a photo naming no camera was received, not shot
min_source_short_side: 1080 # Drop smaller clips unless they carry camera EXIF; 0 or more pixels
max_source_video_seconds: 300 # Exclude longer source videos on Immich metadata, before download (0 disables)
# Album source
max_album_assets: 10000 # Most assets read from one album, per media type (min 1)
# Downloads
download_workers: 3 # Parallel download clients for video and thumbnail prefetching (1-8)
source_prepare_workers: auto # Sources prepared at once (1-4); auto: reserve 1 GiB, then 3 GiB per worker; 1-2, capped by available CPUs
# Duration sizing
optimal_clip_duration: 5.0 # Expected seconds per clip when a trip or album sizes its own duration (2-15s)
# Live Photos (iPhone 3s video clips)
include_live_photos: true # Include Live Photo clips (ON by default)
live_photo_merge_window_seconds: 10.0 # Max gap to group as burst (1-60s)
live_photo_min_clip_seconds: 3.5 # Below this a burst plays its kept picture's own clip (0-30s)
The filename globs exclude matching source media before preparation. Replacing the list replaces
all defaults; [] disables that filter. exclude_stills_without_camera_exif excludes photos
whose EXIF names no camera; videos are exempt. Set it to false for exported originals that lost
camera metadata. These filters apply before selection, so a model cannot bring excluded media back.
Any Live Photo cluster of two or more within the merge window is treated as a burst; the count is not configurable. Where a clip is cut, and how long it runs, is the editor's decision per carrier.
max_album_assets applies per media type, so the default reads up to 10,000 videos and 10,000
photos from one album. Immich returns newest first, so a bigger album is truncated to its most
recent assets, with a warning naming it.
Speech boundaries
advanced:
speech:
enabled: true # Move video cuts out of detected speech
vad_threshold: 0.25 # Voice probability threshold (0.1-0.9)
min_silence_ms: 200 # Pause that separates utterances (50-2000ms)
The bundled FireRedVAD model runs locally with the editorial or editorial-cuda extra.
It measures retained videos and Live Photo companions, then maps speech onto the stitched
timeline. The editor fits the resulting intervals before rendering; an uninterrupted
utterance may cost more time or cause a clip to be left out. This detects voice activity,
not sentence meaning.
The same model also scores a clip's own sound for music or singing, which is how the soundtrack
knows to step aside instead of just ducking under it. Without the ONNX runtime, or with
advanced.speech.enabled: false, that detection never runs and the soundtrack never steps aside.
Generation defaults
defaults:
scale_mode: "blur" # blur | fit (black bars); used when --scale-mode is not given
transition: "smart" # cut, crossfade, smart, none (used when --transition is left on smart)
transition_duration: 0.5 # 0-2 seconds
sharing: "family" # just-us | family | shareable; used when --sharing is not given
add_date: true # caption each clip with its date; --no-add-date turns it off for one film
add_place: true # caption each clip with its place; --no-add-place turns it off for one film
add_date and add_place are the one rule for captions on every surface: the web render panel
starts from them, generate and runs render use them when neither --add-date nor
--no-add-date is given, and automation follows them. Set one to false to turn that caption off
everywhere, automation included. Privacy mode never shows a place, whatever add_place says.
sharing is who a film is for when the run doesn't say (generate --sharing, or Who will watch
it in the web brief). just-us is the household: a private moment a caption names, like a bath,
plays too. family keeps those out. shareable plays only what nothing held back. The rules:
Sharing levels.
Target duration and orientation are per run (--duration, --orientation, or the UI), with the
memory type preset supplying the default duration; there is no config default for either. The
target covers the finished video, title and ending cards included, and the encoder lands near it
rather than exactly on it. There is no backfill: if the stories the editor funded do not fill the
budget, the run reports the shortfall.
Output
output:
directory: "~/Videos/Memories"
format: "mp4" # mp4 or mov
resolution: "1080p" # 720p, 1080p, 4k
codec: h264 # h264 (default), h265 (HDR-capable), prores
codec_policy: prefer_hardware # prefer_hardware (default) or strict
hdr_mode: auto # auto, sdr, hdr
quality: "balanced" # high, balanced, fast (shorthand for CRF presets)
crf: null # unset = derived from quality; 0-51 overrides (lower = better)
min_free_space_gb: 5.0 # warn below this on the output/cache volumes; 0.5-1000
CRF is the image-quality authority. quality is only a shorthand used when crf is omitted; an
explicit crf wins. The number is on libx265's CRF scale, and every other encoder is
calibrated against it: each backend gets whatever setting reproduces the same picture, measured by
SSIM, rather than the same integer. Lower CRF means higher quality everywhere. The measured table
per encoder is on
the hardware overview.
The presets are points on that curve, measured on 1080p60 film:
quality | reference CRF | SSIM | software bitrate | per minute |
|---|---|---|---|---|
high | 18 | 0.99169 | 4.6 Mbps | ~35 MB |
balanced (default) | 24 | 0.98451 | 1.6 Mbps | ~12 MB |
fast | 24 | 0.98451 | 1.6 Mbps | ~12 MB, encoded as fast as the backend can |
There is no tier below balanced: around SSIM 0.980 gradients start to band. fast keeps the
balanced picture and buys its speed from the encoder effort preset instead, overriding
hardware.encoder_preset. Choose balanced, high or fast for a new configuration.
codec_policy decides what happens when the machine has no hardware encoder for the codec you
asked for but does have one for the other. prefer_hardware (the default) switches codec and says
so in the log and the run record, which on a chip like Intel Gemini Lake (H.264 encode entrypoint,
no HEVC one) is the difference between a film finishing and the CPU doing all of it. strict
always honours output.codec and accepts the CPU cost. The switch never applies to ProRes, and
never to an HDR output.
Containers and codecs pair up: mp4 and mov with h264, h265 or prores, and ProRes requires
MOV. generate --format accepts only mp4, h265, and prores: they select H.264/MP4,
H.265/MP4, and ProRes/MOV respectively. Internal and UI overrides also represent h264_mov and
h265_mov, but h264_mov and h265_mov are not CLI choices.
hdr_mode: auto preserves detected HLG or PQ sources when codec: h265 is selected, converting
SDR clips, photos and title screens into the chosen HDR transfer before blending. H.264 is always
SDR: with codec: h264, auto tone-maps detected HDR sources and logs the reason. Use
hdr_mode: sdr when SDR is intentional, or hdr_mode: hdr with H.265 to force an HDR output from
SDR sources.
min_free_space_gb is a preflight, not a cap: it runs before a film starts rendering, on both
output.directory and cache.directory. Below the threshold the run logs a warning naming the
volume and its free space; below what the film itself needs (estimated from target duration and
quality) it stops before writing anything. A run that uploads to Immich has nothing else to do
here: its local film is removed once the upload is confirmed. A run that keeps files locally is
the one this protects. See health, logs and caches.
Photos
photos:
enabled: true # Include photos in memories
duration: 4.0 # Seconds per photo clip (1-10)
burst_window_seconds: 300 # Near-identical photos this close apart are one burst (0-3600)
burst_hash_threshold: 8 # Hash bits two photos may differ by and still be one burst (0-64)
duration is how long the cut holds a still. With no model reader, an empty scene gets half a
second less and the film's first and last stills half a second more; a full film keeps fewer
stills rather than shaving them below that. generate --photo-duration sets it for one run.
The animation per photo (Ken Burns, face pan) is picked from the photo's content and is not
configurable. The bands around a photo that does not fill the canvas follow defaults.scale_mode,
as they do for video.
Burst de-duplication keeps only the best-scored frame of a run of near-identical photos, so fifteen
shots of the same jump do not become fifteen clips. burst_window_seconds: 0 all but turns it off:
photos sharing an identical timestamp still group.
Hardware acceleration
advanced:
hardware:
enabled: true # false = CPU encoding, no GPU probing at all
backend: "auto" # auto, none, nvidia, apple, vaapi, qsv
encoder_preset: "balanced" # fast, balanced, quality
gpu_decode: true # Hardware video decoding
auto detects the backend (NVIDIA NVENC → Apple VideoToolbox → Intel QSV → VAAPI, first hit wins).
hardware.enabled: false or hardware.backend: none forces software rendering. On multi-GPU Linux hosts pick the card with
CUDA_VISIBLE_DEVICES / NVIDIA_VISIBLE_DEVICES.
Naming a backend probes that one and nothing else, which is for measuring rather than for running.
A named backend that cannot encode here logs a warning and falls back to software. backend covers
the video render and Live Photo burst merge. Clip extraction detects its own backend and does
not use hardware.backend.
encoder_preset controls encoder speed and effort; it does not replace output.crf. On Apple,
fast enables VideoToolbox's speed-priority mode while balanced and quality leave it disabled.
Audio and music
Background music uses a bundled track by default. Those tracks ship in Docker and the music
extra; a base pip/uv install needs that extra for the bundled fallback. Generated music needs ace_step.enabled or musicgen.enabled. With both on, ACE-Step generates
and MusicGen is the fallback generator and the stem separator used for ducking; with MusicGen off,
stems come from a local Demucs install if there is one. Per run, --music PATH uses your own file
and --no-music skips music. Music volume is per run too (--music-volume); the ducking and the
2 s / 3 s fades are fixed.
audio:
local_music_dir: "~/Music/Memories" # Library scanned by `immich-memories music search`
max_regenerations: 2 # Extra auto-mode takes when the first is flagged (0-3)
music_block_seconds: 120 # Longest single take before auto mode chains distinct takes (30-300)
max_music_blocks: 3 # Distinct takes to chain for a longer video (1-6)
advanced:
musicgen:
enabled: false # Use a MusicGen API server
base_url: "http://localhost:8000"
api_key: ""
timeout_seconds: 10800 # 3 hours (60-18000)
num_versions: 3 # Versions generated for selection (1-5)
hemisphere: "north" # north or south, for seasonal prompts
ace_step:
enabled: false # Use ACE-Step (remote server or local library)
mode: "api" # api (remote REST server) or lib (local, requires Python 3.12)
api_url: "http://localhost:8000"
api_key: "" # Bearer token for a protected ACE-Step server (api mode)
model_variant: "turbo" # Default 2B; use acestep-v15-xl-turbo for the 4B production profile
lm_model_size: "1.7B" # Default planner; use 4B with the XL production profile
use_lm: false
cpu_offload: true # Local CUDA: move models back to CPU between phases
num_versions: 3 # 1-5
hemisphere: "north"
timeout_seconds: 3600 # 60-18000
advanced.ace_step.cpu_offload defaults to true for local generation. CUDA moves models back
to CPU between phases to reduce VRAM use; false keeps them resident when enough VRAM is
available. The existing Apple Silicon runtime ignores this option, and API requests stay
unchanged. Host and container memory guards still apply. The environment override is
IMMICH_MEMORIES_ACE_STEP__CPU_OFFLOAD=false.
audio.local_music_dir only feeds the immich-memories music helper commands; generation never
picks music from it on its own: pass the file with --music.
audio.max_regenerations bounds auto mode's reaction to a generated track the cheap quality gate
flags as a repetitive "tic-tac". The first take is scored; if it is flagged, auto mode generates
up to that many more takes and keeps the best-scored one. It never drops music, so a run ends with
a track even when every take is flagged.
A video longer than audio.music_block_seconds is not one long generation. Auto mode generates up
to audio.max_music_blocks distinct same-caption takes and joins them with crossfades, then loops
the sequence to fill the remaining length. One long take reads as a metronomic ramble, and one
short phrase on repeat is its own kind of monotony; a chain of a few distinct takes is neither.
Text model and optional vision captions
One llm section serves the selection reader, titles, music mood, special-day scans and explicitly enabled LLM captions. Any OpenAI-compatible or Anthropic-compatible endpoint
works: mlx-vlm, oMLX, Ollama, vLLM, Groq, OpenAI, Claude, z.ai.
advanced:
llm:
enabled: false # false: no LLM calls; true: local or API, chosen by base_url
provider: "openai-compatible" # openai-compatible | openai | zai | anthropic | ollama
base_url: "" # blank: owned locally or hosted provider preset; URL: use that API
model: "gemma-4-E4B-it-Q4_0" # default local Gemma, another GGUF path, or API model name
local_server: "llama-server" # local executable on PATH or its full path
local_mmproj: "" # custom GGUF vision projector; default Gemma has a pinned pair
local_context: 32768 # bounded local context; affects memory use; at least 1024 tokens
api_key: "" # optional, only for cloud APIs
timeout_seconds: 300 # increase for slow local models (10-3600)
preflight_timeout_seconds: 10 # reader availability checks only (greater than 0, at most 3600)
send_image_detail: true # off: APIs whose strict schema rejects image_url.detail
always_reasons: false # true: the endpoint thinks on every call, asked or not
thinking: "disabled" # disabled | low | high | max | auto
reader_concurrency: # independent reader jobs; unset reads it from base_url
batch: "off" # off | auto: provider-supported batching of independent prompts
batch_min_requests: 8 # fewest independent prompts in a stage worth queueing; 2-10000
batch_max_wait_minutes: 60 # then ask whatever the batch has not answered in real time; 1-1440 minutes
# thinking_params: # what the switch looks like on your server
# chat_template_kwargs: # (default: the Qwen dialect, vLLM/mlx)
# enable_thinking: true
# no_thinking_params: # how to say "don't reason" to that server
# chat_template_kwargs: # (default: the Qwen dialect, vLLM/mlx)
# enable_thinking: false
The selection reader receives text: dates, people and place names, and descriptions. A hosted reader sends that text to its provider. Images are sent to the LLM only with explicit editorial.preparation.caption_provider: llm; otherwise captions use their separately configured service. See Privacy.
enabled: false stops requests even when a model and URL remain configured. Set it to true to use a reader. The default model is Gemma.
An empty base_url with enabled: true and openai-compatible or ollama runs llama.cpp
on native Linux or macOS. The Docker and Kubernetes app images need an external server.
Set a URL to use your own API server. openai, anthropic and zai fill a blank URL with their vendor endpoint and select its adapter and reasoning dialect; under zai the URL also picks the adapter (a
.../api/anthropic base takes the Messages route). Which dialect goes where, what thinking does
on each host, how thinking_params and no_thinking_params differ, and what batching pays are all
in Providers and dialects, with links to dated measurements.
thinking has five settings. disabled never asks for reasoning. low, high and max run the
model in reasoning mode for title generation. Bulk reader, music mood and special-day calls use the non-thinking settings. auto sends no reasoning field and takes the host's default, which is where to
start on a host whose dialect you do not know. true and false still parse, as high and
disabled. Reasoning is refused alongside images whatever this is set to: reasoning over several
pictures is a measured runaway. Measured on the live endpoint, a thinking call ran 30-134 s where
the same model answered in 4-7 s without it, and needs a 4000-token ceiling to finish.
The openai preset sends reasoning_effort: none for gpt-5.6-luna, gpt-6-luna and their dated snapshots
on non-thinking calls. Older GPT-5 models keep minimal. An explicit setting wins over the preset.
A provider that rejects a reasoning value reports that error; it does not silently remove the
control and fall back to default reasoning. Only rejection of the parameter itself permits that
fallback.
always_reasons covers a reasoning model that bills its private thinking inside max_tokens, so
the budget the reader asked for its answer is the budget the thinking spends first. Measured on one
hosted API with the same 17 KB monthly read: 245 thinking tokens on the lightest model, 4,126 and
6,256 on two others and 13,469 on the heaviest, all returning HTTP 200 and an empty answer at the
reader's 4,000-token ask. Such calls now ask for the cheapest reasoning the host sells and add
16,384 tokens of room on top of the caller's cap, so the cap keeps meaning what it says about the
answer. The room is a ceiling, not a bill. It is learned from the first reply that reports reasoning
tokens and remembered per server and model; set always_reasons: true to spare that first call,
which otherwise comes back empty.
reader_concurrency limits independent reader jobs in flight (1 to 16). Independent episode-evidence
packs, the period account's calendar-month pages, event inventories and worthiness blocks
can overlap. Pages within an event and later dependent picks remain sequential. Scheduling
preserves prompt text, judgment keys and source ordering; batch delivery is configured separately.
The owned reader always serializes requests. For external servers, unset concurrency is read
from base_url: 1 for loopback, a private IP or a bare service name, 4 for a dotted DNS name
or public IP. A dotted LAN name still gets the hosted policy; set reader_concurrency: 1
for a server that cannot handle overlap. See
Reader concurrency. A provider that
answers 429 pauses every reader in the run, each waiting a slightly different span.
send_image_detail sends OpenAI's optional image_url.detail field. Set it to false for strict
vision schemas that reject anything beyond image_url.url; the zai preset already does.
A provider's dialect can be declared up front instead of negotiated:
advanced:
llm:
max_tokens_param: max_completion_tokens # example; the default is max_tokens
drop_params: [temperature] # example; the default is []
extra_params: {} # fields merged into every call
max_tokens_param and drop_params are read only on the OpenAI dialect, where the query layer
otherwise learns them from the provider's 400s and remembers the answer per server and model.
extra_params also applies on Ollama, where anything under options (num_ctx, num_predict) is
merged into Ollama's own options block rather than replacing it, and a num_predict you set there
wins over the reasoning room the run would otherwise compute.
Two settings shape what a prose request asks for:
advanced:
llm:
structured_output: null # default: select the mode by request type
repetition_penalty: 1.0 # default: sent to a server on your own machine or network, and to Ollama
structured_output sends the JSON shape the request's parser reads as response_format
json_schema, or as Ollama's format. Left unset, it selects the mode for each request:
- Free-text questions, titles and period accounts keep their schemas on both local and hosted endpoints.
- Episode readings on your machine or private network use the shape written in the prompt. oMLX's constrained decoder can stall on that nested schema after an empty array, spending the completion budget before reaching the next required key. Hosted episode readings keep the schema.
Both modes can run against the same endpoint in one process. Set true or false to override
this selection for an endpoint. A provider that refuses schema mode is retried in its compatible
mode, and the run remembers the refusal. Answer-cache identities include the policy and the
request's effective schema, so changing modes cannot reuse an answer from the old policy.
repetition_penalty is sent because local servers default to 1.1 (oMLX, Ollama), which penalises
the repeated keys every JSON answer needs. It is never sent to a public host; a server of your own
that refuses it is asked again without it. Set it to null to leave the server's default.
One llm section supplies all model calls: titles, the selection reader, music mood, special days and optional LLM captions. enabled: true is required; naming a model or URL does not enable it. Move any former separate title-model settings into advanced.llm; that block is removed without a compatibility fallback.
Triage heads
advanced:
triage:
encoder: ~/.immich-memories/models/triage/dinov2-small.onnx # DINOv2-small ONNX export (88 MB)
encoder_url: https://github.com/... # where `models fetch` downloads that export from
provider: auto # ONNX Runtime provider for the encoder: auto, cpu, cuda, coreml
provider: auto takes CUDA where that provider is present and CPU everywhere else. CoreML is opt-in. The provider is operational and does not enter the encoder key, so switching providers does not invalidate matching facts.
Editorial preparation uses triage.encoder with the public eight-head bundle from
editorial.preparation.head_bundle, and checks its digest on load. Missing required head facts
stop selection.
Editorial planner
advanced:
editorial:
reader: rules # derived from the product tier; not an independent choice
thin_model_layer: true # the model polishes a rules draft; false makes it plan the film
strict_sharing: true # anything a head or exposure flag marked stays out of shared films
annotation_database: "" # compatibility field; leave blank for the default cache directory
laya_audience: false # derived: off for Basic, on for GPU and Full
detectors_enabled: false # derived: Marqo and Docling off for Basic, on for GPU and Full
# Apple silicon defaults below; elsewhere the ONNX archive and threshold 0.185 are used.
laya_checkpoint: "~/.immich-memories/models/laya/laya-audience-a79ad9fa.tar"
laya_checkpoint_url: "https://github.com/sam-dumont/immich-memories/releases/download/models-v2/laya-audience-a79ad9fa.tar"
laya_audience_threshold: 0.186 # Laya's hold probability at or above which a carrier is held; 0-1
description_model: "smolvlm2-500m-base-public@envelope-v3-compact"
pixel_producer_key: "pixel-facts-v1" # exact producer of pixel facts and thresholds # gitleaks:allow
head_versions: # exact producer version selected for each annotation head
activity: public-v1
children: public-v1
doc_docling: det-v2
frame_kind: public-v1
location: public-v1
nsfw_marqo: det-v3
people: public-v1
screen: public-v1-strict
uncovered_person: public-v1
venue: oi-v3
preparation:
tier: no_captions # internal producer mode, derived from the product tier
caption_provider: smolvlm # llm explicitly opts into images sent to advanced.llm; higher cost
caption_base_url: http://localhost:8092/v1
caption_artifact_id: "" # optional artifact/revision label; existing captions stay banked; at most 512 characters
caption_api_key: "" # bearer token for a caption server that requires one
caption_timeout_seconds: 90 # positive seconds
caption_concurrency: 1 # 1-16; raise it for a captioner on a GPU
batch_size: 32 # 1-256
head_bundle: "" # packaged public eight-head bundle
detector_python: "" # current Python interpreter
detector_cache_dir: "" # normal Hugging Face Hub cache
marqo_onnx: ~/.immich-memories/models/detectors/nsfw-marqo-384.onnx # digest-pinned sensitive-content export
marqo_onnx_url: https://github.com/sam-dumont/immich-memories/releases/download/models-v1/nsfw-marqo-384-924658f1.onnx
allow_model_downloads: false
people:
seat_min_pictures: 20 # a close family member on this many pictures with no shot gets one; at least 1
seat_min_share: 0.05 # ...or on this share of the period's pictures; greater than 0, at most 1
big_story_density: 2.0 # a story is big only at this multiple of the median photographed day...; greater than 0
big_story_family_share: 0.3 # ...with at least this share of its pictures showing close family; 0-1
In YAML, place this section under advanced:.
people is how the people registry's close family (partner, child, parent) reach the selection. A close
family member on at least seat_min_pictures of the period's pictures, or seat_min_share of them,
who is in none of the film's shots gets one seat: see
the family seat. A story without three favourites is floored
to major only when it is both dense (big_story_density times the period's median photographed
day, in pictures per day) and mostly close family (big_story_family_share of its pictures); both
defaults were measured on real months, see
editing without a language model.
Docling uses det-v2, because ONNX layout optimization mislabels documents on the Celeron J4125.
Saved doc_docling: det-v1 settings upgrade on load, and the next run recomputes that head's facts
and refreshes dependent readings; other head facts stay reusable.
nsfw_marqo uses det-v3, which reads a video on up to eight frames across its length rather than
on the single early frame Immich serves as its preview, and keeps the strongest answer. A still is
read exactly as det-v2 read it, so saved det-v1 and det-v2 settings upgrade on load, a still
keeps its banked det-v2 answer as its det-v3 one, and the next run recomputes that head for
videos only. Everything else it banked stays reusable.
frame_kind, screen and uncovered_person were distilled from a typed picture reader onto the
same encoder the other five heads run on, so a library prepared before they existed is owed only
their three rows and keeps everything else it banked. public-v1-strict says screen ships at one
strict band, baked into its coefficients because the bundle schema holds no threshold: at that band
it answered yes on 37 of 3,564 photographs and every one of them was a screen. Each of the three
only adds to a rule another producer already answered, and none of them can clear anything.
The top-level tier sets reader: rules on basic and gpu, model on full.
Rules cover the ten standard memory products, including albums and recurring dates,
from dates, places, favourites, people metadata and whatever preparation facts exist. They reuse
the normal allocation, spacing, audience and timing checks, omit a thesis, keep unsampled Live
Photos as stills, and write no semantic model banks. Saved plans identify the producer as
rules-v1. A custom --start/--end range is cut the same way as a month or a year.
The thin model layer
thin_model_layer decides what a model install spends its calls on. On (the default), the cut is
built by the rules reader, with no model at all, and the model is then asked one closed question
over the finished film: which of these shots adds nothing to it? What the answer and the gates
leave open is refilled from the same stories, and nothing else moves.
It needs a period the library holds an account of, which cataloguing writes:
immich-memories prepare --overviews banks one per calendar month, and a cut of a whole month or
year that finds none writes its own from the readings it has just paid for. Any other single
window (a season, a trip) is its own period and writes its own account the same way. Films over several windows also use the rules draft and refinement by default. false makes the model plan
the whole film even when an account exists.
strict_sharing keeps any picture a detector head or an exposure flag marked out of a shareable
film. Only your own clearance on the picture can lift a detector or exposure hold.
It is also what lets Basic without captions cut a shareable film at all: with it on, a picture every
active head or detector read as clean and nothing flagged is share; with it off, Basic clears
nothing. Just-us
and family films are unchanged. Turning it off does not let a caption clear an exposure flag.
laya_audience answers the sharing question with a local Laya model:
tier: gpu and tier: full turn it on, and immich-memories models fetch downloads it.
Apple Silicon uses laya-mlx; Linux, Windows and Intel Macs use the portable ONNX checkpoint with the appropriate editorial runtime.
| Platform default | Checkpoint under ~/.immich-memories/models/laya/ | Hold threshold |
|---|---|---|
| Apple Silicon | laya-audience-a79ad9fa.tar | 0.186 |
| Other platforms (ONNX) | laya-audience-onnx-90420ef3.tar.gz | 0.185 |
| A missing checkpoint or runtime | ||
| is reported, and the run continues with the conservative rules fallback. | ||
| It reads the compact caption and adds holds; detector and rule holds still apply and | ||
| are never lifted. It works with the rules reader as well as the prose reader. A captioned shot | ||
| without a Laya answer stays held to the family. Sharing never calls an LLM, including exposure | ||
checks and missing-answer fallbacks. The former thin_batched_audience option is removed. | ||
| See Add a reader. |
Preparation tiers
Preparation follows the product tier. Basic acquires pixel facts and the eight shared-DINO heads.
GPU and Full add Marqo, Docling, captions and Laya. detectors_enabled is tier-owned; a saved
value does not override the tier. head_versions keeps all producer identities, while the active
view filters out Marqo and Docling on Basic. Banked facts remain available when you switch back.
Saved review decisions and permanent holds persist across tier changes.
A film starts with a caption-free rules draft, retaining the resolved tier's detector policy, then
captions and checks only selected shots and actual
replacement candidates. prepare remains the explicit job for a larger scope. Changing tier
does not erase banked captions or other producer facts.
The sharing level decides what a caption reading refuses. A caption describing a shirtless baby or a parent holding a newborn in hospital does not, by itself, cause an activity-based hold. Other detector and selection rules still apply. Bath time, breastfeeding, a nappy change and intimate hygiene play only in a just-us film. Graphic medical procedures, sexual content, exposed adult changing and identifying records never play.
Preparation fills missing descriptions, public heads, detectors and pixel measurements in the annotation database, and skips provider calls where the facts are complete. Missing previews or providers stop selection with an explicit incomplete result. See Add captions for the caption endpoint requirements and Inference on a GPU box for the model artifacts.
Inference service
advanced:
inference:
facts_base_url: "" # blank: the heads and detectors run in the app process
timeout_seconds: 60 # one picture, one request; the service answers all producers at once; greater than 0, at most 300 seconds
facts_concurrency: 8 # how many of those requests are in flight at once (1-32)
producers: [heads, nsfw_marqo, doc_docling] # what the service answers for; the rest stay local
fallback_to_local: true # when the service cannot be reached, run the in-process producers
Point facts_base_url at a running inference service
(http://immich-memories-inference:8092 in the Compose profile;
http://inference:8092 in Kubernetes) and prepare and generate send each picture's
preview there once and bank what comes back. The row is the same row the in-process producers
write: same head, version, label and encoder key, because the service runs the application's own
producers and the key is computed over the model artifact, never over where it ran. Change the
provider or the host and nothing is re-derived.
producers narrows what is offloaded. [heads] sends the DINOv2 encoder and the eight context heads
and keeps Marqo and Docling in the app when the tier enables them. Offloading does not
enable producers disabled by the tier.
facts_concurrency is how many pictures are in the air at once. One at a time, measured on a
cluster against a T1000, costs 0.69 s a picture whatever the card is doing, because almost all of
it is the round trip: 3,709 pictures took 42.7 minutes, and a 13,552-picture year would have taken
2.6 hours. Answers are banked in the order the pictures were asked for, so raising this re-derives
nothing. Raise it until the service is the slow half; the ceiling is 32, and the service's own
REQUEST_THREADS decides how many it can answer at once.
When the service does not answer, the failure is recorded against the endpoint in the preparation
report, and with fallback_to_local: true the in-process producers take over for the pictures
still missing facts (which needs the model files from models fetch on the app box). With it off,
configuration loading fails with Cannot verify the required inference service's compute when
the service is unreachable. This can stop any command before it starts, including preflight;
restore the service or turn local fallback on.
Free-text requests
advanced:
free_text:
wordnet: ~/.immich-memories/models/wordnet/wordnet.zip # WordNet 3.0 (11 MB)
wordnet_url: https://raw.githubusercontent.com/... # where `models fetch` gets it
A film asked for in a sentence (experimental, being built) looks the request's words up in
WordNet: is "cat" a thing, is "park" a place, is a town's name also an ordinary word. The corpus is
a model file like the others: models fetch downloads it from a fixed commit of nltk_data and
checks its SHA-256, and a run that finds it missing stops and says to run models fetch. Nothing
fetches it while a film is being made.
Title screens
title_screens:
enabled: true # Opening title, month dividers and ending screen
title_duration: 3.5 # seconds (1-10)
month_divider_duration: 2.0 # seconds (1-5)
ending_duration: 7.0 # seconds (2-15)
map_move_min_seconds: 6.0 # trip map to a nearby place, 2 s still hold included (3-15)
map_move_max_seconds: 8.0 # trip map to a far place, 2 s still hold included (3-15)
locale: "auto" # en fr nl de es it pt-BR pt-PT pl sv ru ja zh-Hans ko, or auto
style_mode: "auto" # auto (mood-based), random, or a named preset
# Named presets: modern_warm, elegant_minimal, vintage_charm, playful_bright, soft_romantic
fade_color: "white" # white or black at the opening and closing
animated_background: true # Animated titles and smooth map flights
show_month_dividers: true # When the video spans several months (all-or-none)
month_divider_threshold: 2 # Min clips in a month to show its divider (1-10)
use_first_name_only: true # "Riley" instead of "Riley Smith" in titles
fade_color chooses the opening and closing fade. animated_background controls movement;
the colour palette and custom fonts are not configurable today.
animated_background: false keeps the gradient still (no rotation, colour pulse or vignette pulse)
and uses three-view map journeys, which is what preset: fast selects. The
immich-memories titles command exposes more of the look as flags for previewing.
Trip detection
trips:
homebase_latitude: 0.0
homebase_longitude: 0.0
min_distance_km: 50 # at least 1 km
min_duration_days: 2 # at least 1 day
max_gap_days: 2 # at least 1 day
Outside calls
Every switch here is off. A run with both off reaches your Immich server, the endpoints you wrote down yourself, and nothing else.
network:
geocoding: false # nominatim.openstreetmap.org
geocoding_url: "" # a self-hosted Nominatim instead of the public one
map_tiles: false # server.arcgisonline.com (World Imagery)
| Key | What it sends | What you get |
|---|---|---|
geocoding | each trip cluster's centroid, and the coordinates of the pictures in the film's window, rounded to about a kilometre, once per place (answers are kept in the store) | the city or village where Immich names a district or neighbouring town, on every name you read (story titles, captions, location cards, map stops, the report), trip names from the map instead of from EXIF, and place names in the film's language |
geocoding_url | the same requests, to this host instead (http://nominatim.lan:8080); empty means nominatim.openstreetmap.org. Only read with geocoding: true | your own Nominatim, nothing sent outside |
map_tiles | tile coordinates covering the trip area and your home base | the trip fly-over, the static trip map, and a satellite background behind location cards |
Fonts are not a switch: a render never downloads one, and titles fonts --install is the one
step that does (Fonts). preflight prints a row
for each switch you turn on, naming the host it will contact.
Cache
cache:
directory: "~/.immich-memories/cache"
database: "~/.immich-memories/cache.db"
video_cache_enabled: true # Cache downloaded videos locally
video_cache_max_size_gb: 10.0 # Max disk usage for video cache (1-500 GB)
video_cache_max_age_days: 7 # Auto-delete cached videos older than this (1-365)
thumbnail_cache_max_size_mb: 10000.0 # Max disk for Immich previews (50 MB-100 GB)
Tight on disk: lower video_cache_max_size_gb, or turn it off with video_cache_enabled: false.
Size the thumbnail cache by your library
thumbnail_cache_max_size_mb is the one cache budget that scales with the library. Every candidate asset in a memory's scope gets an Immich preview fetched and read back several times. Measured on a real library, one preview is about 315 KB, so:
budget in MB ≈ 0.35 × (assets a memory's scope can reach)
A scope of ten thousand candidates wants about 3.4 GB; the 0.35 leaves a little headroom over the measured 0.315 MB. The 10 GB default holds roughly 31,000 previews, which covers three scopes that size.
If the run's working set does not fit, nothing is lost mid-run: previews still in use are never deleted and the cache overflows the limit instead. The next run reclaims them, so the next overlapping memory re-downloads every preview. You get one WARNING per run saying how far over you are. Raise it rather than ignoring it.
The video cache is not library-sized: it holds the originals being assembled, tens of files per run however big your library is. An old config that still sets preview_cache_max_size_mb loads with a warning; it has no effect.
Store database
Where the store lives: owner decisions, the people registry, model answers, run history, automation
state, the special-days catalogue and the settings you edit in the UI.
cache.database only names a pre-store cache.db for the one-time import, and the directory
the run lock files sit in.
database:
url: "sqlite:///~/.immich-memories/store.db" # or postgresql://user:${PGPASSWORD}@host/db
schema: "immich_memories" # PostgreSQL only: the schema holding every table; non-empty
import_from: "" # where the one-time import of pre-store files looks;
# blank = ~/.immich-memories
IMMICH_MEMORIES_DATABASE_URL, IMMICH_MEMORIES_DATABASE_SCHEMA and IMMICH_MEMORIES_IMPORT_FROM
beat the file. All three are read
before the store opens, so the UI can never change them. SQLite on local disk is the default and
fits a single-host install; point url at PostgreSQL 14+ (no extensions) to share a server,
including Immich's own, in a schema of its own. A SQLite file on NFS, SMB or CIFS is refused:
see Environment variables.
Server (UI)
advanced:
server:
host: "127.0.0.1" # Explicit local bind; YAML 0.0.0.0 is ignored (see below)
port: 8080 # Listen port (1-65535)
enable_demo_mode: false # Offer the Demo mode (blur) switch in the web top bar
secure_cookies: false # Mark the session cookie Secure (turn on behind an HTTPS reverse proxy)
trigger_token: "" # Shared secret for POST /api/trigger. Empty, and with auth
# off, the trigger API is not served at all
allow_unauthenticated_lan: false # Listen beyond localhost with auth disabled:
# anyone reaching the port can use the UI and the
# Immich library behind it
allowed_hosts: [] # Hosts answered beyond localhost. Auth off: any other
# Host gets 421. Auth on: empty answers any host, set
# answers only these (and localhost, auth.public_url)
music_upload_quota_mb: 1024 # Room for uploaded soundtracks; the oldest go past it
allowed_hosts is how an unauthenticated LAN install, or a caller that uses another name (a
Kubernetes Service, a NAS hostname), is let in. See
Allowed hosts.
host and port also have CLI flags: immich-memories ui --host 127.0.0.1 --port 9090.
A YAML server.host: 0.0.0.0 is ignored: older versions saved that value automatically.
Without auth or an explicit override, native installs bind to 127.0.0.1; enabling auth
normally permits all interfaces. For deliberate exposure without auth, use ui --host 0.0.0.0,
IMMICH_MEMORIES_SERVER__HOST=0.0.0.0, or allow_unauthenticated_lan: true on a restricted network.
The Docker image already passes --host 0.0.0.0; its host port mapping controls reachability.
Unauthenticated LAN requests also need the hostname in server.allowed_hosts.
See Network and security before changing it.
trigger_token turns on the HTTP trigger: one POST that runs whatever auto run would have
decided, so an Immich workflow (or a cron, or a phone shortcut) can start a memory. See
Trigger from Immich or anything else.
Keep it out of config.yaml with IMMICH_MEMORIES_SERVER__TRIGGER_TOKEN: server does not expand
a ${VAR} reference, so writing one here stores the six literal characters as your token. Either
way the value is redacted from /health, the config viewer and the logs, from config load onwards.
Saving settings does not write config.yaml. Set a LAN bind deliberately, with authentication or allow_unauthenticated_lan: true; otherwise the secure default is localhost.
Upload to Immich
upload:
enabled: false
album_name: null # Created if missing, reused if exists
An uploaded memory is filed on the day of its last picture, in the timezone most of its pictures share, so it lands in your timeline where the memory ends instead of on the day it was rendered. The render day is what you get when no picture in the cut carries a usable time.
Automation
Controls what immich-memories auto suggest and auto run detect and generate. See Automate it for the commands. In YAML, place this section under advanced:.
advanced:
automation:
enabled: false # run the daily auto-run decision inside the web UI process (Docker)
daily_at: "09:00" # HH:MM, local time of that process (container TZ)
cooldown_hours: 24 # min hours between auto-generated memories (1-168)
max_delivery_attempts: 5 # give up on an Immich upload after this many failures (1-50)
upload_to_immich: false # auto-upload results
album_name: null # target album for uploads
detect_monthly: true # monthly highlights candidates
detect_yearly: true # year-in-review candidates
detect_trips: true # GPS trip detection (needs homebase coords)
detect_person_spotlight: true # per-person highlight candidates
detect_activity_burst: true # unusually active months
burst_threshold: 2.0 # multiplier above rolling average to trigger burst; 1-10
special_days_per_year: 6 # days a year discover-days keeps without a model, strongest first; 1-100
accounts: [] # accounts automation reads, as generate --accounts does: list
# primary too, e.g. ["primary", "partner"] (default: primary alone)
detect_groups: true # propose last year's film for each saved people group
# (people group add) whose people have pictures
accounts and saved groups are the automation side of
a second Immich account: the daily scan reads the listed accounts the
same way a manual --accounts run does, one /users/me check per account, and a read that fails
on any of them fails that day's discovery instead of proposing a film from half a household. The
list is exactly what is read, so keep primary in it: ["partner"] alone leaves the primary
library, and trips, out.
Authentication
Protects the web UI. See the Authentication guide for provider-specific setup (OIDC examples, header proxy config, etc.).
advanced:
auth:
enabled: false
provider: basic # basic, oidc, or header
session_ttl_hours: 24 # 1-720
public_url: "" # e.g. https://memories.example.com -- the URL users use to reach you
# on. Pins the OIDC redirect_uri and enables callback-origin
# validation; without it no origin check is performed
# Basic auth
username: ""
password: "" # Supports ${ENV_VAR} expansion
# OIDC / SSO
issuer_url: "" # Auto-discovers via /.well-known/openid-configuration; supports ${ENV_VAR}
client_id: "" # Supports ${ENV_VAR} expansion
client_secret: "" # Supports ${ENV_VAR} expansion; empty for public clients
scope: "openid email profile"
allowed_emails: [] # OIDC: addresses that may sign in. Empty = anyone the IdP
allowed_domains: [] # authenticates. Domains match exactly: example.com does
# not admit sub.example.com
auto_launch: false # Skip login page, redirect straight to IdP
button_text: "Sign in with SSO"
# Trusted header (reverse proxy)
user_header: "Remote-User"
email_header: "Remote-Email"
trusted_proxies: [] # IPs/CIDRs of your proxy. Required for header provider;
# for basic/oidc their X-Forwarded-* headers are trusted
Place under advanced: in your config file (like all Tier 2 sections).
Notifications
Get notified when auto-generation or scheduled jobs complete. Uses Apprise (130+ services: ntfy, Discord, Telegram, Slack, email, webhooks). Apprise ships with the base package, no extra to install. In YAML, place this section under advanced:.
advanced:
notifications:
enabled: false
urls: [] # No destinations by default; add your own Apprise URLs
on_success: true # notify on successful generation
on_failure: true # notify on failed generation
attach_thumbnail: false # opt in; attachments cost bandwidth/provider quota
cooldown_hours: 24 # pause normal attempts after a delivery failure (1-168)
For example, set advanced.notifications.urls: ["ntfy://your-ntfy-host/private-topic"] and
advanced.notifications.enabled: true. Replace the host and topic with your own destination.
Delivery failures are stored as sanitized health state. Normal success and failure
notifications pause during the cooldown instead of hammering a quota-limited provider.
auto test-notification always bypasses the cooldown and a successful test clears it.
Provider URLs, credentials, and response bodies are never included in health output.
Test your config: immich-memories auto test-notification