Skip to main content

Generated audio runtime

Start with Generate music for the supported backend choices and local installation. This reference covers model budgets, runtime repairs and generated-track handling.

Memory and disk​

lib mode checks free memory against the weights the profile has to keep resident, and refuses with a named shortfall rather than letting macOS kill the process mid-render. A refusal is an ordinary backend failure: MusicGen is next, then a bundled track.

On Linux, the budget follows the process's container limit and any tighter parent limit, capped by the host's available memory. Clean, inactive, unmapped file cache can be reclaimed; live allocations and shared, mapped, dirty or pinned pages remain occupied.

ProfileResident weights it needs freeOn disk
XL (4B) with the 4B plannerabout 29 GBabout 28 GB
XL (4B), use_lm: falseabout 21 GBabout 20 GB
2B with the 1.7B plannerabout 11 GBabout 9 GB
2B with the 0.6B plannerabout 9 GBabout 7 GB
2B, use_lm: falseabout 7 GBabout 6 GB

Check this machine with immich-memories capabilities --test-music. It tries the configured profile and the smaller 2B profiles when their weight budgets fit the memory available now. The 0.6B name refers to the planner; the audio generator is still 2B. A reader that stays loaded in oMLX still uses unified memory while idle. Unloading it can make a smaller music profile fit; the command reports a refusal separately from a generation failure.

The app-owned reader releases its model process after selection, before titles and rendering, and before local music or stem separation. The app also closes its selection-owned Laya model and clears unused local model buffers. Later text work can reopen the owned reader, so the stages can use the same memory at different times. An external reader needs its own unload policy. On Apple Silicon's turbo path, inactive Torch weights are parked before native diffusion; the native decoder is released before VAE decoding.

immich-memories capabilities --verify-local

This checks the configured owned reader and local audio with installed weights and synthetic inputs. A verified result covers that smoke check, not every track length or a full film. Use --test-music instead to try fitting smaller profiles; it may download their weights.

Per file, under ~/.cache/ace-step/checkpoints/: the 2B models about 4.5 GB each, XL-turbo about 19 GB, the planners 1.2, 3.4 and 7.8 GB (0.6B, 1.7B, 4B), the shared VAE and embedding about 1.4 GB. Demucs' htdemucs is about 80 MB under ~/.cache/torch/hub/. Old checkpoints are never removed for you.

Weight budgets are not peak-memory guarantees: generation also allocates working buffers. The MLX buffer cache is disabled on Macs with at most 16 GiB physical RAM and capped at 4 GiB on larger Macs. The DiT uses bf16 by default; IMMICH_MEMORIES_ACESTEP_MLX_DIT_FP32=1 selects fp32 and increases its memory use. An external reader can keep memory occupied while idle; unload it through that server’s controls when local music cannot fit.

Local runtime and repairs​

make install-acestep installs ACE-Step v0.1.8 and Demucs into a sibling .venv-acestep, because ACE-Step's Transformers pin wants an older Hugging Face library than the editor. make check-local-audio generates 15 seconds, splits all four stems and fails loudly if any of it didn't happen locally, so a remote server or a bundled track can't pass it. The default check uses the 2B model with its 0.6B planner: about 9 GB resident weights and 7 GB disk. It does not test the 7 GB planner-free profile recommended in the music guide. Every clone and every git worktree needs its own run, since the environment sits next to the checkout. Rerun the installer after moving the checkout or changing the app version; it also repairs operator torchvision::nms does not exist. A bare uv sync can remove Demucs from the editor's environment, and the installer puts it back.

The isolated stack uses patched PyTorch 2.13, TorchAudio 2.11's stable ABI, and TorchVision 0.28; Linux uses CUDA 12.6 wheels; NVIDIA driver 560.35.05 or newer is the recommended baseline for that update series (NVIDIA driver table). CUDA 12 minor compatibility allows older drivers from 525.60.13 with feature and PTX/JIT limits; it does not guarantee every audio kernel works. The CUDA inference container uses 12.8.1; follow its newer recommended driver baseline. The Mac wheels require macOS 14 or newer. This tested stack overrides ACE-Step v0.1.8's older Linux package pins, so its upstream dependency metadata still reports that mismatch. The app disables PyTorch native JIT kernels in its own audio child before PyTorch imports, keeping the existing eager path free of a compiler requirement. A custom direct library process needs TORCH_DISABLE_NATIVE_JIT=1 before Python starts; setting it after importing PyTorch is too late. Local CUDA generation defaults to advanced.ace_step.cpu_offload: true: inactive models move back to CPU between phases to reduce VRAM use. Set it to false only when the card has room to keep them resident. API mode and Apple Silicon ignore this CUDA setting; the host and container memory guard still applies.

Automatic local ACE-Step and Demucs selection checks that a CUDA kernel actually runs and synchronizes. A detected GPU whose installed PyTorch build cannot execute it falls back to CPU. ACE-Step still checks available memory before loading weights.

In a container​

A separate ACE-Step container serves the editor with mode: api and api_url pointing at it; keep its model cache on a volume. On an NVIDIA box that is the way to go. On a Mac, a container can't reach Metal, so run ACE-Step natively and point the app in Docker at http://host.docker.internal:8000.

Use a build containing the model-loading cleanup fix, which keeps a shared card from retaining VRAM after an out-of-memory error:

ghcr.io/sam-dumont/ace-step-1.5@sha256:45530623b81fa48fd8e4c8398d7d4df4c4ff8b16931edfc3c5cf903dee6dfc24

MusicGen​

Meta's MusicGen, through a remote server only: text-to-music, and Demucs stem separation on its /separate endpoint. With ACE-Step enabled it is the fallback; alone, it generates.

advanced:
musicgen:
enabled: true
base_url: "http://musicgen-server:8000"
timeout_seconds: 10800
num_versions: 3

Generated audio is decoded before mastering or stem separation. Empty, unreadable, non-finite (NaN/Inf), and silent tracks (peak at or below -80 dBFS) count as failed generations. The next enabled generator is tried; if all fail, automatic music uses the bundled library and reports the substitution.

What generation adds to the mix​

  • Tempo fits the photos. In a film with photos, the tempo is nudged so a photo lasts a whole number of beats, within the genre's range and 15 % of the mood's tempo. Videos are never re-timed.
  • A stuck loop is re-rolled. Each take is checked for a metronomic, repetitive grid, and a flagged one is replaced by up to audio.max_regenerations (2) more takes, keeping the best. Music is never dropped for it.
  • Long films chain takes. Past audio.music_block_seconds (120), up to audio.max_music_blocks (3) distinct takes are crossfaded and looped, rather than one long generation. ACE-Step bounds each take to the block length while keeping all scene moods in its prompt, even when the scenes describe a longer film. Each take's near-silent lead-in or tail is trimmed first, so a crossfade seam never lands on a silent stretch.
  • Four stems. The track is split with Demucs: vocals duck most under the clips' sound, drums keep their rhythm. Local Demucs uses Metal on Apple Silicon (immich-memories[demucs] alone); the configured inference service handles separation over HTTP, with local fallback controlled by advanced.inference.fallback_to_local. An explicitly enabled MusicGen server's /separate keeps priority.