Skip to main content

Config Reference

Every key with its built-in default. Set one in ~/.immich-memories/config.yaml, as an environment variable, or from the web UI's settings page (saved to the database). Environment beats the file, the file beats the database, the database beats these defaults; immich-memories config show says which one set each key (where a setting comes from).

Config tiers

Tier 2 sections (analysis, hardware, llm, musicgen, ace_step, server, auth, automation, notifications, triage, editorial, inference, free_text, speech) go under an advanced: key in the file:

advanced:
analysis:
max_album_assets: 5000
hardware:
encoder_preset: "quality"

Both placements are read and merge key by key: a top-level value wins for the same key, while other keys under advanced: remain set. Everything else stays top level. Unknown keys inside a section are ignored; unknown top-level keys and invalid values fail validation at startup.

Tier​

One resolved tier controls both preparation and selection. Leave it automatic for normal use.

tier: auto # auto | basic | gpu | full
tierModelsReaderNeeds
basiceight shared-DINO CPU heads, no captionsrulesmodels fetch
gpuheads, Marqo, Docling, captions and Layarulesa caption server, models fetch
fulleverything in gpuan LLM polishes the rules draft and writes the proseenable advanced.llm; owned local model by default, or an explicit base_url and server model

auto is the default. A healthy inference service reporting CUDA, or a local CUDA or MLX/Metal runtime, selects gpu. An explicitly enabled LLM alongside that capability selects full. Without GPU inference capability, selection stays on basic and reports what is missing. A renderer's GPU does not establish inference capability. The runtime check loads no model weights and sends no pictures; preflight and acquisition still check the actual producers.

editorial.reader, editorial.preparation.tier and editorial.laya_audience are derived from the product tier. These values are derived rather than separate user controls. Save omits these derived settings and keeps automatic resolution automatic when the file moves to another host. An explicit basic, gpu or full pins a tier for a controlled comparison; it does not install or start its services. full requires an enabled reader: the app-owned local model or a configured API.

Configured LLM titles and music mood work on every tier. Basic and GPU still select with rules, and sharing never asks the prose LLM. Captions use their own configured service; a text LLM is not an automatic caption fallback. A missing Laya checkpoint or runtime is reported and uses the conservative rules fallback; that is a degraded run, not a verified GPU/full comparison.

Env: IMMICH_MEMORIES_TIER=auto. From a source checkout, uv run python scripts/tier_settings.py prints what each tier runs with. That helper script is not part of a pip/uv install; immich-memories capabilities reports the resolved setup there.

Preset​

One top-level switch that fills several knobs at once. fast is the CPU-only / NAS profile. Anything you set yourself wins over the preset.

preset: null # null | fast

fast sets, unless you set them yourself: output.resolution: 1080p, output.codec: h264, output.quality: fast, hardware.encoder_preset: fast and title_screens.animated_background: false (static title backgrounds and three-view maps). Music generation is already off by default and stays wherever you put it.

Env: IMMICH_MEMORIES_PRESET=fast. One-off on the CLI: immich-memories --preset fast generate … (root option, before the subcommand).

Settings saved in the UI or CLI go to the database. Environment variables and config.yaml still win; the app refuses to save a database value that they would override. Bootstrap keys never go to the database: database.* (read before the store opens), and auth.* and server.* (they decide who can reach the app). Set them in the environment or config.yaml and restart. A saved value cannot contain a ${VAR} reference.

Immich connection​

Immich Memories supports Immich v2 and v3. Automatic runtime detection is the default:

immich:
url: "https://photos.example.com"
api_key: "${IMMICH_API_KEY}"
api_version: auto # auto | v2 | v3

Keep api_version on auto for normal use. The client detects and caches the server major for each runtime client; you do not choose it for each generation. Explicit v2 or v3 is a manual troubleshooting escape hatch for a proxy or unusual deployment that prevents correct detection. An override forces that API contract.

Run the read-only immich-memories config test to check credentials and see the resolved API contract without generating or uploading a memory.

Extra accounts​

People can upload to separate Immich accounts. Put additional accounts under accounts, by name. The top-level url and api_key stay the primary account, and the primary is the only account a film is ever uploaded to.

immich:
url: "https://photos.example.com"
api_key: "${IMMICH_API_KEY}"
native_sharing: false # opt into verified native person identities (3.2+; 3.3 experimental)
accounts: {} # name -> url, api_key, api_version (default: none)
# accounts:
# partner:
# url: "https://photos.example.com"
# api_key: "${PARTNER_IMMICH_API_KEY}"
# api_version: auto

A name is lowercase letters and digits joined by single underscores (partner, grandma_2), and primary is taken: it means the top-level account. The name is what a person alias bound to that account records. Configuring an account adds nothing to your films: a run reads only the accounts it selects, and every selected account has to answer /users/me with its key before anything is read. generate --accounts primary,partner reads the named accounts into one film on the CLI, and automation.accounts does the same for the daily scan. config test and preflight check each one, one line per account, with no key printed: A second Immich account.

The key is a secret like the primary one: redacted from logs and issue reports, and sealed with IMMICH_MEMORIES_SECRET_KEY when saved to the database (the whole immich.accounts map is one encrypted row). From the environment, name the account in the variable: IMMICH_MEMORIES_IMMICH__ACCOUNTS__PARTNER__API_KEY.

Render worker​

The CLI and web UI can send an already selected film to a trusted render worker. Blank worker_base_url renders on the app's machine.

render:
worker_base_url: ""
worker_token: "" # Or ${RENDER_WORKER_TOKEN}
allow_insecure_http: false # Explicitly accept a non-loopback cleartext HTTP worker
timeout_seconds: 3600 # Wait for rendering and download; maximum 86400; greater than 0, at most 86400 seconds
fallback_to_local: false # Set true to allow local rendering after a worker failure

Use the same app version on both machines. The worker receives the selected assets, exact cuts, Live source material, titles, locations, audio markers and the Immich API key so it can download the sources directly. Configure a worker you trust, reachable over your private network or HTTPS. The handoff request carries that Immich key, so a non-loopback http:// worker URL is refused until allow_insecure_http: true says you meant it; loopback addresses and HTTPS need no opt-in. Preflight names the transport before the first render request. Requests require the worker token and do not follow redirects.

The worker accepts H.264 or H.265 MP4. When a worker is configured, MOV and ProRes render locally only with fallback_to_local: true; with false, the run fails. Orientation only sets the canvas; it does not change the selection. Speech detection and cut selection run before handoff. Music and Immich upload finish on the app after it checks the returned film.

See worker deployment for Docker Compose and Kubernetes examples.

Video analysis​

advanced:
analysis:
# Media the camera roll did not shoot (see Configuration → Footage the
# camera roll did not shoot). Setting the list replaces it; [] turns it off.
exclude_filename_patterns: # case-insensitive globs on the source filename
- "RingVideo_*"
- "RPReplay_Final*"
- "Screen Recording *"
- "Screenshot*"
- "img-*-wa[0-9][0-9][0-9][0-9]*"
- "vid-*-wa[0-9][0-9][0-9][0-9]*"
exclude_stills_without_camera_exif: true # a photo naming no camera was received, not shot
min_source_short_side: 1080 # Drop smaller clips unless they carry camera EXIF; 0 or more pixels
max_source_video_seconds: 300 # Exclude longer source videos on Immich metadata, before download (0 disables)

# Album source
max_album_assets: 10000 # Most assets read from one album, per media type (min 1)

# Downloads
download_workers: 3 # Parallel download clients for video and thumbnail prefetching (1-8)
source_prepare_workers: auto # Sources prepared at once (1-4); auto: reserve 1 GiB, then 3 GiB per worker; 1-2, capped by available CPUs

# Duration sizing
optimal_clip_duration: 5.0 # Expected seconds per clip when a trip or album sizes its own duration (2-15s)

# Live Photos (iPhone 3s video clips)
include_live_photos: true # Include Live Photo clips (ON by default)
live_photo_merge_window_seconds: 10.0 # Max gap to group as burst (1-60s)
live_photo_min_clip_seconds: 3.5 # Below this a burst plays its kept picture's own clip (0-30s)

The filename globs exclude matching source media before preparation. Replacing the list replaces all defaults; [] disables that filter. exclude_stills_without_camera_exif excludes photos whose EXIF names no camera; videos are exempt. Set it to false for exported originals that lost camera metadata. These filters apply before selection, so a model cannot bring excluded media back.

Any Live Photo cluster of two or more within the merge window is treated as a burst; the count is not configurable. Where a clip is cut, and how long it runs, is the editor's decision per carrier.

max_album_assets applies per media type, so the default reads up to 10,000 videos and 10,000 photos from one album. Immich returns newest first, so a bigger album is truncated to its most recent assets, with a warning naming it.

Speech boundaries​

advanced:
speech:
enabled: true # Move video cuts out of detected speech
vad_threshold: 0.25 # Voice probability threshold (0.1-0.9)
min_silence_ms: 200 # Pause that separates utterances (50-2000ms)

The bundled FireRedVAD model runs locally with the editorial or editorial-cuda extra. It measures retained videos and Live Photo companions, then maps speech onto the stitched timeline. The editor fits the resulting intervals before rendering; an uninterrupted utterance may cost more time or cause a clip to be left out. This detects voice activity, not sentence meaning.

The same model also scores a clip's own sound for music or singing, which is how the soundtrack knows to step aside instead of just ducking under it. Without the ONNX runtime, or with advanced.speech.enabled: false, that detection never runs and the soundtrack never steps aside.

Generation defaults​

defaults:
scale_mode: "blur" # blur | fit (black bars); used when --scale-mode is not given
transition: "smart" # cut, crossfade, smart, none (used when --transition is left on smart)
transition_duration: 0.5 # 0-2 seconds
sharing: "family" # just-us | family | shareable; used when --sharing is not given
add_date: true # caption each clip with its date; --no-add-date turns it off for one film
add_place: true # caption each clip with its place; --no-add-place turns it off for one film

add_date and add_place are the one rule for captions on every surface: the web render panel starts from them, generate and runs render use them when neither --add-date nor --no-add-date is given, and automation follows them. Set one to false to turn that caption off everywhere, automation included. Privacy mode never shows a place, whatever add_place says.

sharing is who a film is for when the run doesn't say (generate --sharing, or Who will watch it in the web brief). just-us is the household: a private moment a caption names, like a bath, plays too. family keeps those out. shareable plays only what nothing held back. The rules: Sharing levels.

Target duration and orientation are per run (--duration, --orientation, or the UI), with the memory type preset supplying the default duration; there is no config default for either. The target covers the finished video, title and ending cards included, and the encoder lands near it rather than exactly on it. There is no backfill: if the stories the editor funded do not fill the budget, the run reports the shortfall.

Output​

output:
directory: "~/Videos/Memories"
format: "mp4" # mp4 or mov
resolution: "1080p" # 720p, 1080p, 4k
codec: h264 # h264 (default), h265 (HDR-capable), prores
codec_policy: prefer_hardware # prefer_hardware (default) or strict
hdr_mode: auto # auto, sdr, hdr
quality: "balanced" # high, balanced, fast (shorthand for CRF presets)
crf: null # unset = derived from quality; 0-51 overrides (lower = better)
min_free_space_gb: 5.0 # warn below this on the output/cache volumes; 0.5-1000

CRF is the image-quality authority. quality is only a shorthand used when crf is omitted; an explicit crf wins. The number is on libx265's CRF scale, and every other encoder is calibrated against it: each backend gets whatever setting reproduces the same picture, measured by SSIM, rather than the same integer. Lower CRF means higher quality everywhere. The measured table per encoder is on the hardware overview.

The presets are points on that curve, measured on 1080p60 film:

qualityreference CRFSSIMsoftware bitrateper minute
high180.991694.6 Mbps~35 MB
balanced (default)240.984511.6 Mbps~12 MB
fast240.984511.6 Mbps~12 MB, encoded as fast as the backend can

There is no tier below balanced: around SSIM 0.980 gradients start to band. fast keeps the balanced picture and buys its speed from the encoder effort preset instead, overriding hardware.encoder_preset. Choose balanced, high or fast for a new configuration.

codec_policy decides what happens when the machine has no hardware encoder for the codec you asked for but does have one for the other. prefer_hardware (the default) switches codec and says so in the log and the run record, which on a chip like Intel Gemini Lake (H.264 encode entrypoint, no HEVC one) is the difference between a film finishing and the CPU doing all of it. strict always honours output.codec and accepts the CPU cost. The switch never applies to ProRes, and never to an HDR output.

Containers and codecs pair up: mp4 and mov with h264, h265 or prores, and ProRes requires MOV. generate --format accepts only mp4, h265, and prores: they select H.264/MP4, H.265/MP4, and ProRes/MOV respectively. Internal and UI overrides also represent h264_mov and h265_mov, but h264_mov and h265_mov are not CLI choices.

hdr_mode: auto preserves detected HLG or PQ sources when codec: h265 is selected, converting SDR clips, photos and title screens into the chosen HDR transfer before blending. H.264 is always SDR: with codec: h264, auto tone-maps detected HDR sources and logs the reason. Use hdr_mode: sdr when SDR is intentional, or hdr_mode: hdr with H.265 to force an HDR output from SDR sources.

min_free_space_gb is a preflight, not a cap: it runs before a film starts rendering, on both output.directory and cache.directory. Below the threshold the run logs a warning naming the volume and its free space; below what the film itself needs (estimated from target duration and quality) it stops before writing anything. A run that uploads to Immich has nothing else to do here: its local film is removed once the upload is confirmed. A run that keeps files locally is the one this protects. See health, logs and caches.

Photos​

photos:
enabled: true # Include photos in memories
duration: 4.0 # Seconds per photo clip (1-10)
burst_window_seconds: 300 # Near-identical photos this close apart are one burst (0-3600)
burst_hash_threshold: 8 # Hash bits two photos may differ by and still be one burst (0-64)

duration is how long the cut holds a still. With no model reader, an empty scene gets half a second less and the film's first and last stills half a second more; a full film keeps fewer stills rather than shaving them below that. generate --photo-duration sets it for one run.

The animation per photo (Ken Burns, face pan) is picked from the photo's content and is not configurable. The bands around a photo that does not fill the canvas follow defaults.scale_mode, as they do for video.

Burst de-duplication keeps only the best-scored frame of a run of near-identical photos, so fifteen shots of the same jump do not become fifteen clips. burst_window_seconds: 0 all but turns it off: photos sharing an identical timestamp still group.

Hardware acceleration​

advanced:
hardware:
enabled: true # false = CPU encoding, no GPU probing at all
backend: "auto" # auto, none, nvidia, apple, vaapi, qsv
encoder_preset: "balanced" # fast, balanced, quality
gpu_decode: true # Hardware video decoding

auto detects the backend (NVIDIA NVENC → Apple VideoToolbox → Intel QSV → VAAPI, first hit wins). hardware.enabled: false or hardware.backend: none forces software rendering. On multi-GPU Linux hosts pick the card with CUDA_VISIBLE_DEVICES / NVIDIA_VISIBLE_DEVICES.

Naming a backend probes that one and nothing else, which is for measuring rather than for running. A named backend that cannot encode here logs a warning and falls back to software. backend covers the video render and Live Photo burst merge. Clip extraction detects its own backend and does not use hardware.backend.

encoder_preset controls encoder speed and effort; it does not replace output.crf. On Apple, fast enables VideoToolbox's speed-priority mode while balanced and quality leave it disabled.

Audio and music​

Background music uses a bundled track by default. Those tracks ship in Docker and the music extra; a base pip/uv install needs that extra for the bundled fallback. Generated music needs ace_step.enabled or musicgen.enabled. With both on, ACE-Step generates and MusicGen is the fallback generator and the stem separator used for ducking; with MusicGen off, stems come from a local Demucs install if there is one. Per run, --music PATH uses your own file and --no-music skips music. Music volume is per run too (--music-volume); the ducking and the 2 s / 3 s fades are fixed.

audio:
local_music_dir: "~/Music/Memories" # Library scanned by `immich-memories music search`
max_regenerations: 2 # Extra auto-mode takes when the first is flagged (0-3)
music_block_seconds: 120 # Longest single take before auto mode chains distinct takes (30-300)
max_music_blocks: 3 # Distinct takes to chain for a longer video (1-6)
advanced:
musicgen:
enabled: false # Use a MusicGen API server
base_url: "http://localhost:8000"
api_key: ""
timeout_seconds: 10800 # 3 hours (60-18000)
num_versions: 3 # Versions generated for selection (1-5)
hemisphere: "north" # north or south, for seasonal prompts

ace_step:
enabled: false # Use ACE-Step (remote server or local library)
mode: "api" # api (remote REST server) or lib (local, requires Python 3.12)
api_url: "http://localhost:8000"
api_key: "" # Bearer token for a protected ACE-Step server (api mode)
model_variant: "turbo" # Default 2B; use acestep-v15-xl-turbo for the 4B production profile
lm_model_size: "1.7B" # Default planner; use 4B with the XL production profile
use_lm: false
cpu_offload: true # Local CUDA: move models back to CPU between phases
num_versions: 3 # 1-5
hemisphere: "north"
timeout_seconds: 3600 # 60-18000

advanced.ace_step.cpu_offload defaults to true for local generation. CUDA moves models back to CPU between phases to reduce VRAM use; false keeps them resident when enough VRAM is available. The existing Apple Silicon runtime ignores this option, and API requests stay unchanged. Host and container memory guards still apply. The environment override is IMMICH_MEMORIES_ACE_STEP__CPU_OFFLOAD=false.

audio.local_music_dir only feeds the immich-memories music helper commands; generation never picks music from it on its own: pass the file with --music.

audio.max_regenerations bounds auto mode's reaction to a generated track the cheap quality gate flags as a repetitive "tic-tac". The first take is scored; if it is flagged, auto mode generates up to that many more takes and keeps the best-scored one. It never drops music, so a run ends with a track even when every take is flagged.

A video longer than audio.music_block_seconds is not one long generation. Auto mode generates up to audio.max_music_blocks distinct same-caption takes and joins them with crossfades, then loops the sequence to fill the remaining length. One long take reads as a metronomic ramble, and one short phrase on repeat is its own kind of monotony; a chain of a few distinct takes is neither.

Text model and optional vision captions​

One llm section serves the selection reader, titles, music mood, special-day scans and explicitly enabled LLM captions. Any OpenAI-compatible or Anthropic-compatible endpoint works: mlx-vlm, oMLX, Ollama, vLLM, Groq, OpenAI, Claude, z.ai.

advanced:
llm:
enabled: false # false: no LLM calls; true: local or API, chosen by base_url
provider: "openai-compatible" # openai-compatible | openai | zai | anthropic | ollama
base_url: "" # blank: owned locally or hosted provider preset; URL: use that API
model: "gemma-4-E4B-it-Q4_0" # default local Gemma, another GGUF path, or API model name
local_server: "llama-server" # local executable on PATH or its full path
local_mmproj: "" # custom GGUF vision projector; default Gemma has a pinned pair
local_context: 32768 # bounded local context; affects memory use; at least 1024 tokens
api_key: "" # optional, only for cloud APIs
timeout_seconds: 300 # increase for slow local models (10-3600)
preflight_timeout_seconds: 10 # reader availability checks only (greater than 0, at most 3600)
send_image_detail: true # off: APIs whose strict schema rejects image_url.detail
always_reasons: false # true: the endpoint thinks on every call, asked or not
thinking: "disabled" # disabled | low | high | max | auto
reader_concurrency: # independent reader jobs; unset reads it from base_url
batch: "off" # off | auto: provider-supported batching of independent prompts
batch_min_requests: 8 # fewest independent prompts in a stage worth queueing; 2-10000
batch_max_wait_minutes: 60 # then ask whatever the batch has not answered in real time; 1-1440 minutes
# thinking_params: # what the switch looks like on your server
# chat_template_kwargs: # (default: the Qwen dialect, vLLM/mlx)
# enable_thinking: true
# no_thinking_params: # how to say "don't reason" to that server
# chat_template_kwargs: # (default: the Qwen dialect, vLLM/mlx)
# enable_thinking: false

The selection reader receives text: dates, people and place names, and descriptions. A hosted reader sends that text to its provider. Images are sent to the LLM only with explicit editorial.preparation.caption_provider: llm; otherwise captions use their separately configured service. See Privacy.

enabled: false stops requests even when a model and URL remain configured. Set it to true to use a reader. The default model is Gemma.

An empty base_url with enabled: true and openai-compatible or ollama runs llama.cpp on native Linux or macOS. The Docker and Kubernetes app images need an external server. Set a URL to use your own API server. openai, anthropic and zai fill a blank URL with their vendor endpoint and select its adapter and reasoning dialect; under zai the URL also picks the adapter (a .../api/anthropic base takes the Messages route). Which dialect goes where, what thinking does on each host, how thinking_params and no_thinking_params differ, and what batching pays are all in Providers and dialects, with links to dated measurements.

thinking has five settings. disabled never asks for reasoning. low, high and max run the model in reasoning mode for title generation. Bulk reader, music mood and special-day calls use the non-thinking settings. auto sends no reasoning field and takes the host's default, which is where to start on a host whose dialect you do not know. true and false still parse, as high and disabled. Reasoning is refused alongside images whatever this is set to: reasoning over several pictures is a measured runaway. Measured on the live endpoint, a thinking call ran 30-134 s where the same model answered in 4-7 s without it, and needs a 4000-token ceiling to finish.

The openai preset sends reasoning_effort: none for gpt-5.6-luna, gpt-6-luna and their dated snapshots on non-thinking calls. Older GPT-5 models keep minimal. An explicit setting wins over the preset. A provider that rejects a reasoning value reports that error; it does not silently remove the control and fall back to default reasoning. Only rejection of the parameter itself permits that fallback.

always_reasons covers a reasoning model that bills its private thinking inside max_tokens, so the budget the reader asked for its answer is the budget the thinking spends first. Measured on one hosted API with the same 17 KB monthly read: 245 thinking tokens on the lightest model, 4,126 and 6,256 on two others and 13,469 on the heaviest, all returning HTTP 200 and an empty answer at the reader's 4,000-token ask. Such calls now ask for the cheapest reasoning the host sells and add 16,384 tokens of room on top of the caller's cap, so the cap keeps meaning what it says about the answer. The room is a ceiling, not a bill. It is learned from the first reply that reports reasoning tokens and remembered per server and model; set always_reasons: true to spare that first call, which otherwise comes back empty.

reader_concurrency limits independent reader jobs in flight (1 to 16). Independent episode-evidence packs, the period account's calendar-month pages, event inventories and worthiness blocks can overlap. Pages within an event and later dependent picks remain sequential. Scheduling preserves prompt text, judgment keys and source ordering; batch delivery is configured separately.

The owned reader always serializes requests. For external servers, unset concurrency is read from base_url: 1 for loopback, a private IP or a bare service name, 4 for a dotted DNS name or public IP. A dotted LAN name still gets the hosted policy; set reader_concurrency: 1 for a server that cannot handle overlap. See Reader concurrency. A provider that answers 429 pauses every reader in the run, each waiting a slightly different span.

send_image_detail sends OpenAI's optional image_url.detail field. Set it to false for strict vision schemas that reject anything beyond image_url.url; the zai preset already does.

A provider's dialect can be declared up front instead of negotiated:

advanced:
llm:
max_tokens_param: max_completion_tokens # example; the default is max_tokens
drop_params: [temperature] # example; the default is []
extra_params: {} # fields merged into every call

max_tokens_param and drop_params are read only on the OpenAI dialect, where the query layer otherwise learns them from the provider's 400s and remembers the answer per server and model. extra_params also applies on Ollama, where anything under options (num_ctx, num_predict) is merged into Ollama's own options block rather than replacing it, and a num_predict you set there wins over the reasoning room the run would otherwise compute.

Two settings shape what a prose request asks for:

advanced:
llm:
structured_output: null # default: select the mode by request type
repetition_penalty: 1.0 # default: sent to a server on your own machine or network, and to Ollama

structured_output sends the JSON shape the request's parser reads as response_format json_schema, or as Ollama's format. Left unset, it selects the mode for each request:

  • Free-text questions, titles and period accounts keep their schemas on both local and hosted endpoints.
  • Episode readings on your machine or private network use the shape written in the prompt. oMLX's constrained decoder can stall on that nested schema after an empty array, spending the completion budget before reaching the next required key. Hosted episode readings keep the schema.

Both modes can run against the same endpoint in one process. Set true or false to override this selection for an endpoint. A provider that refuses schema mode is retried in its compatible mode, and the run remembers the refusal. Answer-cache identities include the policy and the request's effective schema, so changing modes cannot reuse an answer from the old policy.

repetition_penalty is sent because local servers default to 1.1 (oMLX, Ollama), which penalises the repeated keys every JSON answer needs. It is never sent to a public host; a server of your own that refuses it is asked again without it. Set it to null to leave the server's default.

One llm section supplies all model calls: titles, the selection reader, music mood, special days and optional LLM captions. enabled: true is required; naming a model or URL does not enable it. Move any former separate title-model settings into advanced.llm; that block is removed without a compatibility fallback.

Triage heads​

advanced:
triage:
encoder: ~/.immich-memories/models/triage/dinov2-small.onnx # DINOv2-small ONNX export (88 MB)
encoder_url: https://github.com/... # where `models fetch` downloads that export from
provider: auto # ONNX Runtime provider for the encoder: auto, cpu, cuda, coreml

provider: auto takes CUDA where that provider is present and CPU everywhere else. CoreML is opt-in. The provider is operational and does not enter the encoder key, so switching providers does not invalidate matching facts.

Editorial preparation uses triage.encoder with the public eight-head bundle from editorial.preparation.head_bundle, and checks its digest on load. Missing required head facts stop selection.

Editorial planner​

advanced:
editorial:
reader: rules # derived from the product tier; not an independent choice
thin_model_layer: true # the model polishes a rules draft; false makes it plan the film
strict_sharing: true # anything a head or exposure flag marked stays out of shared films
annotation_database: "" # compatibility field; leave blank for the default cache directory
laya_audience: false # derived: off for Basic, on for GPU and Full
detectors_enabled: false # derived: Marqo and Docling off for Basic, on for GPU and Full
# Apple silicon defaults below; elsewhere the ONNX archive and threshold 0.185 are used.
laya_checkpoint: "~/.immich-memories/models/laya/laya-audience-a79ad9fa.tar"
laya_checkpoint_url: "https://github.com/sam-dumont/immich-memories/releases/download/models-v2/laya-audience-a79ad9fa.tar"
laya_audience_threshold: 0.186 # Laya's hold probability at or above which a carrier is held; 0-1
description_model: "smolvlm2-500m-base-public@envelope-v3-compact"
pixel_producer_key: "pixel-facts-v1" # exact producer of pixel facts and thresholds # gitleaks:allow
head_versions: # exact producer version selected for each annotation head
activity: public-v1
children: public-v1
doc_docling: det-v2
frame_kind: public-v1
location: public-v1
nsfw_marqo: det-v3
people: public-v1
screen: public-v1-strict
uncovered_person: public-v1
venue: oi-v3
preparation:
tier: no_captions # internal producer mode, derived from the product tier
caption_provider: smolvlm # llm explicitly opts into images sent to advanced.llm; higher cost
caption_base_url: http://localhost:8092/v1
caption_artifact_id: "" # optional artifact/revision label; existing captions stay banked; at most 512 characters
caption_api_key: "" # bearer token for a caption server that requires one
caption_timeout_seconds: 90 # positive seconds
caption_concurrency: 1 # 1-16; raise it for a captioner on a GPU
batch_size: 32 # 1-256
head_bundle: "" # packaged public eight-head bundle
detector_python: "" # current Python interpreter
detector_cache_dir: "" # normal Hugging Face Hub cache
marqo_onnx: ~/.immich-memories/models/detectors/nsfw-marqo-384.onnx # digest-pinned sensitive-content export
marqo_onnx_url: https://github.com/sam-dumont/immich-memories/releases/download/models-v1/nsfw-marqo-384-924658f1.onnx
allow_model_downloads: false
people:
seat_min_pictures: 20 # a close family member on this many pictures with no shot gets one; at least 1
seat_min_share: 0.05 # ...or on this share of the period's pictures; greater than 0, at most 1
big_story_density: 2.0 # a story is big only at this multiple of the median photographed day...; greater than 0
big_story_family_share: 0.3 # ...with at least this share of its pictures showing close family; 0-1

In YAML, place this section under advanced:.

people is how the people registry's close family (partner, child, parent) reach the selection. A close family member on at least seat_min_pictures of the period's pictures, or seat_min_share of them, who is in none of the film's shots gets one seat: see the family seat. A story without three favourites is floored to major only when it is both dense (big_story_density times the period's median photographed day, in pictures per day) and mostly close family (big_story_family_share of its pictures); both defaults were measured on real months, see editing without a language model.

Docling uses det-v2, because ONNX layout optimization mislabels documents on the Celeron J4125. Saved doc_docling: det-v1 settings upgrade on load, and the next run recomputes that head's facts and refreshes dependent readings; other head facts stay reusable.

nsfw_marqo uses det-v3, which reads a video on up to eight frames across its length rather than on the single early frame Immich serves as its preview, and keeps the strongest answer. A still is read exactly as det-v2 read it, so saved det-v1 and det-v2 settings upgrade on load, a still keeps its banked det-v2 answer as its det-v3 one, and the next run recomputes that head for videos only. Everything else it banked stays reusable.

frame_kind, screen and uncovered_person were distilled from a typed picture reader onto the same encoder the other five heads run on, so a library prepared before they existed is owed only their three rows and keeps everything else it banked. public-v1-strict says screen ships at one strict band, baked into its coefficients because the bundle schema holds no threshold: at that band it answered yes on 37 of 3,564 photographs and every one of them was a screen. Each of the three only adds to a rule another producer already answered, and none of them can clear anything.

The top-level tier sets reader: rules on basic and gpu, model on full. Rules cover the ten standard memory products, including albums and recurring dates, from dates, places, favourites, people metadata and whatever preparation facts exist. They reuse the normal allocation, spacing, audience and timing checks, omit a thesis, keep unsampled Live Photos as stills, and write no semantic model banks. Saved plans identify the producer as rules-v1. A custom --start/--end range is cut the same way as a month or a year.

The thin model layer​

thin_model_layer decides what a model install spends its calls on. On (the default), the cut is built by the rules reader, with no model at all, and the model is then asked one closed question over the finished film: which of these shots adds nothing to it? What the answer and the gates leave open is refilled from the same stories, and nothing else moves.

It needs a period the library holds an account of, which cataloguing writes: immich-memories prepare --overviews banks one per calendar month, and a cut of a whole month or year that finds none writes its own from the readings it has just paid for. Any other single window (a season, a trip) is its own period and writes its own account the same way. Films over several windows also use the rules draft and refinement by default. false makes the model plan the whole film even when an account exists.

strict_sharing keeps any picture a detector head or an exposure flag marked out of a shareable film. Only your own clearance on the picture can lift a detector or exposure hold. It is also what lets Basic without captions cut a shareable film at all: with it on, a picture every active head or detector read as clean and nothing flagged is share; with it off, Basic clears nothing. Just-us and family films are unchanged. Turning it off does not let a caption clear an exposure flag.

laya_audience answers the sharing question with a local Laya model: tier: gpu and tier: full turn it on, and immich-memories models fetch downloads it. Apple Silicon uses laya-mlx; Linux, Windows and Intel Macs use the portable ONNX checkpoint with the appropriate editorial runtime.

Platform defaultCheckpoint under ~/.immich-memories/models/laya/Hold threshold
Apple Siliconlaya-audience-a79ad9fa.tar0.186
Other platforms (ONNX)laya-audience-onnx-90420ef3.tar.gz0.185
A missing checkpoint or runtime
is reported, and the run continues with the conservative rules fallback.
It reads the compact caption and adds holds; detector and rule holds still apply and
are never lifted. It works with the rules reader as well as the prose reader. A captioned shot
without a Laya answer stays held to the family. Sharing never calls an LLM, including exposure
checks and missing-answer fallbacks. The former thin_batched_audience option is removed.
See Add a reader.

Preparation tiers​

Preparation follows the product tier. Basic acquires pixel facts and the eight shared-DINO heads. GPU and Full add Marqo, Docling, captions and Laya. detectors_enabled is tier-owned; a saved value does not override the tier. head_versions keeps all producer identities, while the active view filters out Marqo and Docling on Basic. Banked facts remain available when you switch back. Saved review decisions and permanent holds persist across tier changes.

A film starts with a caption-free rules draft, retaining the resolved tier's detector policy, then captions and checks only selected shots and actual replacement candidates. prepare remains the explicit job for a larger scope. Changing tier does not erase banked captions or other producer facts.

The sharing level decides what a caption reading refuses. A caption describing a shirtless baby or a parent holding a newborn in hospital does not, by itself, cause an activity-based hold. Other detector and selection rules still apply. Bath time, breastfeeding, a nappy change and intimate hygiene play only in a just-us film. Graphic medical procedures, sexual content, exposed adult changing and identifying records never play.

Preparation fills missing descriptions, public heads, detectors and pixel measurements in the annotation database, and skips provider calls where the facts are complete. Missing previews or providers stop selection with an explicit incomplete result. See Add captions for the caption endpoint requirements and Inference on a GPU box for the model artifacts.

Inference service​

advanced:
inference:
facts_base_url: "" # blank: the heads and detectors run in the app process
timeout_seconds: 60 # one picture, one request; the service answers all producers at once; greater than 0, at most 300 seconds
facts_concurrency: 8 # how many of those requests are in flight at once (1-32)
producers: [heads, nsfw_marqo, doc_docling] # what the service answers for; the rest stay local
fallback_to_local: true # when the service cannot be reached, run the in-process producers

Point facts_base_url at a running inference service (http://immich-memories-inference:8092 in the Compose profile; http://inference:8092 in Kubernetes) and prepare and generate send each picture's preview there once and bank what comes back. The row is the same row the in-process producers write: same head, version, label and encoder key, because the service runs the application's own producers and the key is computed over the model artifact, never over where it ran. Change the provider or the host and nothing is re-derived.

producers narrows what is offloaded. [heads] sends the DINOv2 encoder and the eight context heads and keeps Marqo and Docling in the app when the tier enables them. Offloading does not enable producers disabled by the tier.

facts_concurrency is how many pictures are in the air at once. One at a time, measured on a cluster against a T1000, costs 0.69 s a picture whatever the card is doing, because almost all of it is the round trip: 3,709 pictures took 42.7 minutes, and a 13,552-picture year would have taken 2.6 hours. Answers are banked in the order the pictures were asked for, so raising this re-derives nothing. Raise it until the service is the slow half; the ceiling is 32, and the service's own REQUEST_THREADS decides how many it can answer at once.

When the service does not answer, the failure is recorded against the endpoint in the preparation report, and with fallback_to_local: true the in-process producers take over for the pictures still missing facts (which needs the model files from models fetch on the app box). With it off, configuration loading fails with Cannot verify the required inference service's compute when the service is unreachable. This can stop any command before it starts, including preflight; restore the service or turn local fallback on.

Free-text requests​

advanced:
free_text:
wordnet: ~/.immich-memories/models/wordnet/wordnet.zip # WordNet 3.0 (11 MB)
wordnet_url: https://raw.githubusercontent.com/... # where `models fetch` gets it

A film asked for in a sentence (experimental, being built) looks the request's words up in WordNet: is "cat" a thing, is "park" a place, is a town's name also an ordinary word. The corpus is a model file like the others: models fetch downloads it from a fixed commit of nltk_data and checks its SHA-256, and a run that finds it missing stops and says to run models fetch. Nothing fetches it while a film is being made.

Title screens​

title_screens:
enabled: true # Opening title, month dividers and ending screen
title_duration: 3.5 # seconds (1-10)
month_divider_duration: 2.0 # seconds (1-5)
ending_duration: 7.0 # seconds (2-15)
map_move_min_seconds: 6.0 # trip map to a nearby place, 2 s still hold included (3-15)
map_move_max_seconds: 8.0 # trip map to a far place, 2 s still hold included (3-15)
locale: "auto" # en fr nl de es it pt-BR pt-PT pl sv ru ja zh-Hans ko, or auto
style_mode: "auto" # auto (mood-based), random, or a named preset
# Named presets: modern_warm, elegant_minimal, vintage_charm, playful_bright, soft_romantic
fade_color: "white" # white or black at the opening and closing
animated_background: true # Animated titles and smooth map flights
show_month_dividers: true # When the video spans several months (all-or-none)
month_divider_threshold: 2 # Min clips in a month to show its divider (1-10)
use_first_name_only: true # "Riley" instead of "Riley Smith" in titles

fade_color chooses the opening and closing fade. animated_background controls movement; the colour palette and custom fonts are not configurable today. animated_background: false keeps the gradient still (no rotation, colour pulse or vignette pulse) and uses three-view map journeys, which is what preset: fast selects. The immich-memories titles command exposes more of the look as flags for previewing.

Trip detection​

trips:
homebase_latitude: 0.0
homebase_longitude: 0.0
min_distance_km: 50 # at least 1 km
min_duration_days: 2 # at least 1 day
max_gap_days: 2 # at least 1 day

Outside calls​

Every switch here is off. A run with both off reaches your Immich server, the endpoints you wrote down yourself, and nothing else.

network:
geocoding: false # nominatim.openstreetmap.org
geocoding_url: "" # a self-hosted Nominatim instead of the public one
map_tiles: false # server.arcgisonline.com (World Imagery)
KeyWhat it sendsWhat you get
geocodingeach trip cluster's centroid, and the coordinates of the pictures in the film's window, rounded to about a kilometre, once per place (answers are kept in the store)the city or village where Immich names a district or neighbouring town, on every name you read (story titles, captions, location cards, map stops, the report), trip names from the map instead of from EXIF, and place names in the film's language
geocoding_urlthe same requests, to this host instead (http://nominatim.lan:8080); empty means nominatim.openstreetmap.org. Only read with geocoding: trueyour own Nominatim, nothing sent outside
map_tilestile coordinates covering the trip area and your home basethe trip fly-over, the static trip map, and a satellite background behind location cards

Fonts are not a switch: a render never downloads one, and titles fonts --install is the one step that does (Fonts). preflight prints a row for each switch you turn on, naming the host it will contact.

Cache​

cache:
directory: "~/.immich-memories/cache"
database: "~/.immich-memories/cache.db"
video_cache_enabled: true # Cache downloaded videos locally
video_cache_max_size_gb: 10.0 # Max disk usage for video cache (1-500 GB)
video_cache_max_age_days: 7 # Auto-delete cached videos older than this (1-365)
thumbnail_cache_max_size_mb: 10000.0 # Max disk for Immich previews (50 MB-100 GB)

Tight on disk: lower video_cache_max_size_gb, or turn it off with video_cache_enabled: false.

Size the thumbnail cache by your library​

thumbnail_cache_max_size_mb is the one cache budget that scales with the library. Every candidate asset in a memory's scope gets an Immich preview fetched and read back several times. Measured on a real library, one preview is about 315 KB, so:

budget in MB ≈ 0.35 × (assets a memory's scope can reach)

A scope of ten thousand candidates wants about 3.4 GB; the 0.35 leaves a little headroom over the measured 0.315 MB. The 10 GB default holds roughly 31,000 previews, which covers three scopes that size.

If the run's working set does not fit, nothing is lost mid-run: previews still in use are never deleted and the cache overflows the limit instead. The next run reclaims them, so the next overlapping memory re-downloads every preview. You get one WARNING per run saying how far over you are. Raise it rather than ignoring it.

The video cache is not library-sized: it holds the originals being assembled, tens of files per run however big your library is. An old config that still sets preview_cache_max_size_mb loads with a warning; it has no effect.

Store database​

Where the store lives: owner decisions, the people registry, model answers, run history, automation state, the special-days catalogue and the settings you edit in the UI. cache.database only names a pre-store cache.db for the one-time import, and the directory the run lock files sit in.

database:
url: "sqlite:///~/.immich-memories/store.db" # or postgresql://user:${PGPASSWORD}@host/db
schema: "immich_memories" # PostgreSQL only: the schema holding every table; non-empty
import_from: "" # where the one-time import of pre-store files looks;
# blank = ~/.immich-memories

IMMICH_MEMORIES_DATABASE_URL, IMMICH_MEMORIES_DATABASE_SCHEMA and IMMICH_MEMORIES_IMPORT_FROM beat the file. All three are read before the store opens, so the UI can never change them. SQLite on local disk is the default and fits a single-host install; point url at PostgreSQL 14+ (no extensions) to share a server, including Immich's own, in a schema of its own. A SQLite file on NFS, SMB or CIFS is refused: see Environment variables.

Server (UI)​

advanced:
server:
host: "127.0.0.1" # Explicit local bind; YAML 0.0.0.0 is ignored (see below)
port: 8080 # Listen port (1-65535)
enable_demo_mode: false # Offer the Demo mode (blur) switch in the web top bar
secure_cookies: false # Mark the session cookie Secure (turn on behind an HTTPS reverse proxy)
trigger_token: "" # Shared secret for POST /api/trigger. Empty, and with auth
# off, the trigger API is not served at all
allow_unauthenticated_lan: false # Listen beyond localhost with auth disabled:
# anyone reaching the port can use the UI and the
# Immich library behind it
allowed_hosts: [] # Hosts answered beyond localhost. Auth off: any other
# Host gets 421. Auth on: empty answers any host, set
# answers only these (and localhost, auth.public_url)
music_upload_quota_mb: 1024 # Room for uploaded soundtracks; the oldest go past it

allowed_hosts is how an unauthenticated LAN install, or a caller that uses another name (a Kubernetes Service, a NAS hostname), is let in. See Allowed hosts.

host and port also have CLI flags: immich-memories ui --host 127.0.0.1 --port 9090. A YAML server.host: 0.0.0.0 is ignored: older versions saved that value automatically. Without auth or an explicit override, native installs bind to 127.0.0.1; enabling auth normally permits all interfaces. For deliberate exposure without auth, use ui --host 0.0.0.0, IMMICH_MEMORIES_SERVER__HOST=0.0.0.0, or allow_unauthenticated_lan: true on a restricted network. The Docker image already passes --host 0.0.0.0; its host port mapping controls reachability. Unauthenticated LAN requests also need the hostname in server.allowed_hosts. See Network and security before changing it.

trigger_token turns on the HTTP trigger: one POST that runs whatever auto run would have decided, so an Immich workflow (or a cron, or a phone shortcut) can start a memory. See Trigger from Immich or anything else. Keep it out of config.yaml with IMMICH_MEMORIES_SERVER__TRIGGER_TOKEN: server does not expand a ${VAR} reference, so writing one here stores the six literal characters as your token. Either way the value is redacted from /health, the config viewer and the logs, from config load onwards.

Saving settings does not write config.yaml. Set a LAN bind deliberately, with authentication or allow_unauthenticated_lan: true; otherwise the secure default is localhost.

Upload to Immich​

upload:
enabled: false
album_name: null # Created if missing, reused if exists

An uploaded memory is filed on the day of its last picture, in the timezone most of its pictures share, so it lands in your timeline where the memory ends instead of on the day it was rendered. The render day is what you get when no picture in the cut carries a usable time.

Automation​

Controls what immich-memories auto suggest and auto run detect and generate. See Automate it for the commands. In YAML, place this section under advanced:.

advanced:
automation:
enabled: false # run the daily auto-run decision inside the web UI process (Docker)
daily_at: "09:00" # HH:MM, local time of that process (container TZ)
cooldown_hours: 24 # min hours between auto-generated memories (1-168)
max_delivery_attempts: 5 # give up on an Immich upload after this many failures (1-50)
upload_to_immich: false # auto-upload results
album_name: null # target album for uploads
detect_monthly: true # monthly highlights candidates
detect_yearly: true # year-in-review candidates
detect_trips: true # GPS trip detection (needs homebase coords)
detect_person_spotlight: true # per-person highlight candidates
detect_activity_burst: true # unusually active months
burst_threshold: 2.0 # multiplier above rolling average to trigger burst; 1-10
special_days_per_year: 6 # days a year discover-days keeps without a model, strongest first; 1-100
accounts: [] # accounts automation reads, as generate --accounts does: list
# primary too, e.g. ["primary", "partner"] (default: primary alone)
detect_groups: true # propose last year's film for each saved people group
# (people group add) whose people have pictures

accounts and saved groups are the automation side of a second Immich account: the daily scan reads the listed accounts the same way a manual --accounts run does, one /users/me check per account, and a read that fails on any of them fails that day's discovery instead of proposing a film from half a household. The list is exactly what is read, so keep primary in it: ["partner"] alone leaves the primary library, and trips, out.

Authentication​

Protects the web UI. See the Authentication guide for provider-specific setup (OIDC examples, header proxy config, etc.).

advanced:
auth:
enabled: false
provider: basic # basic, oidc, or header
session_ttl_hours: 24 # 1-720
public_url: "" # e.g. https://memories.example.com -- the URL users use to reach you
# on. Pins the OIDC redirect_uri and enables callback-origin
# validation; without it no origin check is performed

# Basic auth
username: ""
password: "" # Supports ${ENV_VAR} expansion

# OIDC / SSO
issuer_url: "" # Auto-discovers via /.well-known/openid-configuration; supports ${ENV_VAR}
client_id: "" # Supports ${ENV_VAR} expansion
client_secret: "" # Supports ${ENV_VAR} expansion; empty for public clients
scope: "openid email profile"
allowed_emails: [] # OIDC: addresses that may sign in. Empty = anyone the IdP
allowed_domains: [] # authenticates. Domains match exactly: example.com does
# not admit sub.example.com
auto_launch: false # Skip login page, redirect straight to IdP
button_text: "Sign in with SSO"

# Trusted header (reverse proxy)
user_header: "Remote-User"
email_header: "Remote-Email"
trusted_proxies: [] # IPs/CIDRs of your proxy. Required for header provider;
# for basic/oidc their X-Forwarded-* headers are trusted

Place under advanced: in your config file (like all Tier 2 sections).

Notifications​

Get notified when auto-generation or scheduled jobs complete. Uses Apprise (130+ services: ntfy, Discord, Telegram, Slack, email, webhooks). Apprise ships with the base package, no extra to install. In YAML, place this section under advanced:.

advanced:
notifications:
enabled: false
urls: [] # No destinations by default; add your own Apprise URLs
on_success: true # notify on successful generation
on_failure: true # notify on failed generation
attach_thumbnail: false # opt in; attachments cost bandwidth/provider quota
cooldown_hours: 24 # pause normal attempts after a delivery failure (1-168)

For example, set advanced.notifications.urls: ["ntfy://your-ntfy-host/private-topic"] and advanced.notifications.enabled: true. Replace the host and topic with your own destination.

Delivery failures are stored as sanitized health state. Normal success and failure notifications pause during the cooldown instead of hammering a quota-limited provider. auto test-notification always bypasses the cooldown and a successful test clears it. Provider URLs, credentials, and response bodies are never included in health output.

Test your config: immich-memories auto test-notification