Local Stable Audio SFX
Generate game sound effects locally with Stable Audio 3 Small SFX on Apple Silicon using MLX, including WAV masters, MP3 takes and processed ambience loops.
The official Stable Audio 3 MLX runtime
lives at ~/AI/stable-audio-3/optimized/mlx. Its project-local .venv uses Python 3.11 and MLX/Apple Metal.
The selected model is sm-sfx, paired with same-s and the shared T5Gemma text encoder. Music (towns, the GM room and engaged named boss encounters,
ADR-026) uses the optional sm-music bundle through --model sm-music; it shares the same codec
and text encoder.
Set STABLE_AUDIO_HOME to override the runtime directory. No remote generation service is used.
◆Generate
From the Zoen root:
./tools/sfxgen --prompt "heavy sword impact on stone, isolated game sound effect" \
--seconds 1.2 --out audio/source/weapons/sword-impact.wav
The user-local sfxgen command works from any directory. Output paths are relative to the caller's working directory;
absolute paths also work. Running the repository launcher by its absolute path works from elsewhere too.
Parent directories are created. Existing outputs (including any numbered take) are refused unless --overwrite is given.
Prompts pass through unchanged: creature grunts, breaths, roars and chants receive no automatic "no voices" suffix.
| Option | Behavior |
|---|---|
--model sm-sfx|sm-music | sm-sfx (default) for effects and ambience; sm-music for looping music (the Small Music bundle) |
--seconds N | Required final duration, greater than zero and at most 120 seconds; no time stretching |
--takes N | First file uses the supplied name; subsequent files use _2, _3, etc. |
--format wav|mp3 | Inferred from extension, otherwise WAV; an explicit format must agree with the extension |
--seed N | Official deterministic seed control; subsequent takes use N+1, N+2; reproducibility depends on runtime/version |
--negative-prompt TEXT | Passed through with guidance enabled (--cfg 3), because default CFG 1 ignores it |
--normalize | Optional peak normalization to −3 dBFS; do not routinely use on quiet ambience |
--loop | Overlap crossfade for ambience; preserves steady level rather than trimming/fading to silence |
--crossfade N | Loop overlap in seconds (default 0.25; Zoen's music uses 2). Needs --loop; at most a tenth of --seconds |
--play | Play each validated result with macOS afplay |
--overwrite | Explicitly replace existing requested output files |
# Three takes, high-quality MP3 (WAV is the temporary master)
./tools/sfxgen --prompt "fast lightning snap and compact thunder hit" \
--seconds 1.2 --takes 3 --seed 42 --out audio/source/skills/lightning.mp3
# WAV loop master; audition its repeated boundary before shipping.
# Keep "no music" in --negative-prompt: a negation inside the prompt can attract what it excludes.
./tools/sfxgen --prompt "quiet fantasy meadow, soft wind and distant wildlife" \
--negative-prompt "music, melody, singing, voices, speech" \
--seconds 20 --loop --out audio/source/ambience/meadow.wav
# Looping music with the Small Music bundle: a 60 s loop with a 2 s overlap at the seam
./tools/sfxgen --model sm-music --prompt "classic orchestral fantasy town theme, flute melody over strings and harp, 92 BPM, instrumental" \
--negative-prompt "vocals, singing, choir, electronic synthesizer, drum kit" \
--seconds 60 --takes 3 --loop --crossfade 2 --out audio/source/music/town.mp3
# Playback
./tools/sfxgen --prompt "tiny bronze button tick" --seconds 0.3 --play --out /tmp/button.wav
◆Processing and verification
For non-loop effects below 1.5 seconds, generate at least 1.5 seconds of source. Detect active 5 ms blocks then refine onset to the first active sample, using a relative −40 dB threshold with a −60 dBFS floor and a robust baseline; keep up to 10 ms before the first activity (capped at one tenth of the requested duration so very short outputs still contain the attack). Trim inactive leading/trailing blocks, preserve the first attack, crop to the requested duration and pad remaining space with silence. Apply 2 ms anti-click fades and remove DC offset. Cropping can shorten a long tail; audition final effects. No stretching, compressor, or automatic semantic prompt rewrite is applied.
By default gain is only reduced when the peak exceeds −3 dBFS, matching Zoen's SFX source ceiling.
--normalize also raises quieter effects to that ceiling. Ambience is not LUFS-normalized here:
Zoen's separate delivery/mixer pipeline still owns its −16 LUFS target and platform encodings.
The official CLI normally clamps its decoder output when saving 16-bit PCM. A small adapter imports the official
CLI and replaces only its WAV writer with a float WAV writer. Inference and sampling remain official; float masters
preserve peaks above 0 dBFS so gain reduction happens before quantization and clipping. No upstream code is edited.
Loops ask the model for the loop plus its overlap plus 3 s of guard. The model ends every clip with 0–1.5 s of digital silence
(about −95 dBFS, and the length changes with the requested duration), so that silent tail is detected and cut before anything
is built. The body and the post-roll then come only from real sound; the post-roll is mixed into the opening with a linear
overlap crossfade (250 ms by default, --crossfade to change it), and the exact requested duration is kept. A source that is
still too short after the cut is refused. Every loop is then checked: its quieter edge (first or last 100 ms) must not sit more
than 30 dB below the clip's median level, and the result prints edge_dip_db. At the 120-second model limit, blend both edges
toward their shared midpoint instead of requesting unsupported extra duration (and ask for a shorter loop if that source is
silent at its end). This smooths the boundary; it does not guarantee that a generative phrase or moving sound will loop
perceptually. Prefer WAV for loops; MP3 players may introduce gaps.
WAV output is a lossless 24-bit PCM master at the runtime's 44.1 kHz stereo rate (temporary inference masters retain float samples).
MP3 uses FFmpeg/libmp3lame VBR quality 2 and is verified by full decoding. If decoded peaks exceed −3 dBFS,
re-encode from the lossless WAV with measured attenuation (up to three attempts); no lossy intermediate is transcoded.
Temporary WAVs are removed.
Every output is decoded and checked for finite samples, nonzero energy, rate, channels, duration and peak before publication.
Outputs publish atomically; a concurrent writer cannot be silently overwritten without --overwrite.
A later take failing leaves already successful takes in place and returns an error.
◆Zoen audio requests
Audio pipeline and packages/game-data/scripts/gen-audio-requests.mjs specify
audio/source/<category>/<id>.mp3, requested seconds, variations and loop flags. Each request now carries the exact command (command in audio-requests.json and
audio-requests/queue.json); Codex loops the queue with the prompt on the Audio generation page. Take names match the catalog. WAV masters can use the same stem (the catalog also recognizes WAV).
After intentionally generating production assets, record each source/license/date in Licences
and run pnpm data:build.
Final mono/stereo conversion, 48 kHz resampling, Opus/WebM and AAC packaging remain separate delivery steps.
◆Model files, authorization and updates
The installer symlinks weights under models/mlx/ from the Hugging Face cache at ~/.cache/huggingface/hub/:
dit_sm-sfx_f16.npz, dit_sm-music_f16.npz, same_s_decoder_f32.npz, same_s_encoder_f32.npz, t5gemma_f16.npz.
The Small SFX bundle covers effects and ambience; no Medium, CUDA or TensorRT installation is needed. Music uses the optional
Small Music bundle (dit_sm-music_f16.npz, 919 MB, same Hugging Face repository and licence). It was installed on 2026-10-05
with the owner's OK (OPEN-3): ./install.sh -y --python 3.11 --download sm-music in the runtime directory. If it is ever
missing, --model sm-music stops with that exact command.
Generation sets HF_HUB_OFFLINE=1 and HF_DATASETS_OFFLINE=1; missing model files fail locally instead of downloading.
The two bundles are different weights: Stability describes small-music as "music-only" and small-sfx as "sound-effects-only".
Test on 2026-10-05: the SFX model given the Amber Gate and GM-room music prompts returned pitched, key-consistent sound (key-profile
correlation 0.81, where noise scores 0.37) but no beat at the requested tempo (pulse 0.10 and 0.28, where a 92 BPM reference
scores 0.78). Measured, not heard; the two clips are in logs/evidence/SFX-ENV-20261005-01/.
Hugging Face access may require a read token and personal acceptance of Stability AI/Gemma terms.
If the official download reports a gated access error, accept the terms on the named model page and log in using
~/AI/stable-audio-3/optimized/mlx/.venv/bin/hf auth login in your own terminal. Never place a token in this repository,
CLI arguments, docs or logs. Public anonymous access, when offered by the official repository, needs no token.
Use remains subject to the Stability AI Community License and the applicable text-encoder terms;
review current terms for your intended release.
To update runtime code while preserving local changes:
cd ~/AI/stable-audio-3
git status --short
git pull --ff-only
cd optimized/mlx
./install.sh --help
./install.sh -y --python 3.11 --download sm-sfx
uv pip install --python .venv/bin/python imageio-ffmpeg
The installer skips present weights. Before deliberately refreshing weights, preserve the working cache/revision, remove only the four model symlinks listed above, then rerun the official bundle installer and reverify inference. Do not delete the entire shared Hugging Face cache. Rerun the wrapper tests after any update.
◆Troubleshooting
- Missing runtime: check
STABLE_AUDIO_HOME,.venv/bin/python, andsa3. - Missing weights/offline errors: rerun the official bundle install while online, then verify all four symlink targets exist.
- MP3 needs FFmpeg. This Mac uses the ARM64 binary bundled by
imageio-ffmpegin the runtime venv;~/.local/bin/ffmpegexposes it. If absent, install the small wheel with the command above. An existing system FFmpeg is preferred. WAV processing does not require FFmpeg or ffprobe. - Memory errors: close heavy apps; this wrapper keeps the official progressive model freeing enabled.
- Silent or malformed audio: the command fails verification; try a clearer prompt or another seed.
- Gain/headroom: the float-master adapter avoids the official WAV clamp; inspect the attack by ear.
- Confirm GPU availability and the default device:
~/AI/stable-audio-3/optimized/mlx/.venv/bin/python -c \
'import mlx.core as mx; print(mx.metal.is_available(), mx.default_device())'
# Expected: True Device(gpu, 0)
The wrapper also checks this before inference and refuses an unavailable Metal GPU. Run focused checks with the runtime interpreter:
~/AI/stable-audio-3/optimized/mlx/.venv/bin/python -m unittest discover -s tools/audio -p 'test_*.py'
Source: zoen/docs/LOCAL-SFX.md · 1,644 words · edit the Markdown, not this page.
