$ env HF_HUB_OFFLINE=1 /Users/king/AI/stable-audio-3/optimized/mlx/sa3 --prompt Short fantasy sword hitting a metal shield, sharp metallic clang, compact impact, isolated sound effect, no music, no intelligible speech, clean start. --seconds 1.5 --dit sm-sfx --decoder same-s --seed 20261004 --out /Users/king/Documents/Codex/2026-10-04/files-pasted-by-the-user-i/work/raw-test/shield.wav # cwd /Users/king/Desktop/Workspace/SRO/zoen # exit 2 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Examples: ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ tip: the ./sa3 wrapper handles .venv + uv automatically 🎵 Generate audio from a prompt $ ./sa3 --prompt "footsteps on gravel, then a door slamming" \ --dit sm-sfx --decoder same-s --seconds 8 --out sfx.wav # sound-effect generation ▶ Play immediately after generation $ ./sa3 --prompt "ambient drone" --dit sm-sfx --decoder same-s \ --seconds 10 --out drone.wav --play # writes WAV + plays via afplay (Ctrl-C stops both) 🎚️ Audio-to-audio & inpainting (requires an input WAV) $ ./sa3 --prompt "jazz fusion with electric piano" --dit sm-sfx --decoder same-s \ --init-audio funk.wav --init-noise-level 0.7 --out funk_jazz.wav # variation: 0.4-0.8 typical, higher = more change $ ./sa3 --prompt "explosive drum break" --dit sm-sfx --decoder same-s \ --init-audio funk.wav --inpaint-range "4,7" --out funk_drums.wav # regenerate seconds 4-7, keep rest 🎯 Steer with CFG + negative prompts $ ./sa3 --prompt "ambient drone" --cfg 3.0 \ --negative-prompt "drums, vocals, distortion" \ --dit sm-sfx --decoder same-s --out clean_drone.wav # cfg > 1.0 toward prompt, neg pushes away note: bundles not installed: medium, sm-music. Re-run ./install.sh to pick them up, or just use them in ./sa3 — weights auto-download from HuggingFace on first use. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ error: unrecognized arguments: fantasy sword hitting a metal shield, sharp metallic clang, compact impact, isolated sound effect, no music, no intelligible speech, clean start. usage: sa3_mlx.py [-h] [--prompt PROMPT] [--negative-prompt NEGATIVE_PROMPT] [--init-audio INIT_AUDIO] [--inpaint-range INPAINT_RANGE] [--dit {sm-music,sm-sfx,medium}] [--decoder {same-s,same-l}] [--dit-dtype {fp32,fp16}] [--t5gemma-npz T5GEMMA_NPZ] [--lora ADAPTER [KEY=VAL ...]] [--lora-strength LORA_STRENGTH] [--seconds SECONDS] [--steps STEPS] [--seed SEED] [--init-noise-level INIT_NOISE_LEVEL] [--cfg CFG] [--apg APG] [--free-models | --no-free-models] [--out OUT] [--play] SA3 text-to-audio (+ audio-to-audio + inpainting) in pure MLX options: -h, --help show this help message and exit --prompt PROMPT Text prompt describing the audio to generate. Empty string is valid (unconditional generation). If omitted, the script asks interactively via stdin. --negative-prompt NEGATIVE_PROMPT Negative prompt for CFG's unconditional branch. When --cfg=1.0 this flag has no effect (no uncond pass is run). When unset and --cfg ≠ 1.0, the uncond branch uses zero embeddings. --init-audio INIT_AUDIO Path to a WAV file (44.1 kHz, 16-bit PCM, stereo or mono) to use as the starting point. Enables audio-to- audio mode (with --init-noise-level) or inpainting mode (with --inpaint-range). Encoder is loaded automatically. Audio is trimmed or zero-padded to match --seconds. --inpaint-range INPAINT_RANGE Inpainting time range as 'START,END' in seconds (e.g. '5,10'). Requires --init-audio. The model regenerates the masked range while preserving the rest of the input exactly (via paste-back). --dit {sm-music,sm-sfx,medium} DiT model. 'small' = sa3-sm-music (faster, smaller). 'medium' = sa3-medium-ARC (larger, higher quality). If omitted, prompts interactively with arrow-key picker. --decoder {same-s,same-l} Audio decoder. 'same-s' pairs with sa3-small (50 M params). 'same-l' pairs with sa3-medium (426 M params). If omitted, prompts interactively with arrow- key picker. --dit-dtype {fp32,fp16} DiT compute dtype. Default fp16 — validated transparent at ~50-57 dB PSNR vs FP32. Halves DiT memory and speeds sampling ~25%. The decoder always runs FP32 (SAME-S catastrophically cancels at fp16 due to differential attention) and T5Gemma is always fp16. Set to fp32 only if you need bit-exact reproducibility against the PyTorch reference. --t5gemma-npz T5GEMMA_NPZ Path to the bundled T5Gemma FP16 .npz (weights + tokenizer). Default points at models/mlx/t5gemma_f16.npz next to this script; auto- downloaded from HuggingFace if not present. --lora ADAPTER [KEY=VAL ...] A LoRA adapter to apply to the DiT — repeat the flag for several. Each --lora takes the adapter path followed by optional key=value options: strength=S (default --lora-strength) and steps=RANGE, a 1-based inclusive sampling-step range: 2-8, 2- (2..last), -4 (1..4) or 3 (just step 3; default: all steps). Example: --lora myband.safetensors strength=0.8 steps=2- . Adapters covering every step are merged at load; step-gated ones are re-merged in place at the (few) step boundaries — ~80 ms each on medium, no extra memory, every step runs at full speed. The adapter is a .safetensors file (SA3-native train_lora.py / underfit output) or a PEFT adapter directory (with its adapter_config.json). ONLY .safetensors is accepted — a pickle .ckpt/.pt is refused (it would execute code on load). The adapter's base must match --dit. --lora-strength LORA_STRENGTH Default strength for --lora adapters without their own strength= (default 1.0). 0 disables (bit-identical to no LoRA); >1 amplifies. --seconds SECONDS Output audio length in seconds. T_lat (latent positions) is derived as ceil(seconds * 44100 / 4096) — a natural ceil that depends ONLY on --seconds (and --seed via the DiT), never on --decoder. This keeps the sampled latent (and hence the music) identical across decoders for a given prompt/seed. Final WAV is trimmed to exactly --seconds. --steps STEPS Number of pingpong sampling steps. Minimum 1 (single forward pass — fastest, lowest quality). rf_denoiser is distilled for 8 (default — sweet spot). >8 gives diminishing returns and may overshoot. The LogSNR schedule auto-computes for any N: steps=N produces N+1 sigma values from σmax→0. Sample wall time scales ~linearly with steps; quality/coherence improves noticeably from 1→4 and 4→8, less from 8→16. --seed SEED Random seed (any int). Set this for reproducible outputs. If omitted, a random seed is chosen and printed in the final 'done' line. --init-noise-level INIT_NOISE_LEVEL σmax — the schedule's starting noise level (always honored regardless of mode). Valid range: [0.01, ∞). Below 0.01 the script errors because the model is undefined at t≈0 and produces NaN. With --init-audio: 0.4-0.8 is typical for variation; 1.0 = full regeneration (init ignored). Without --init-audio: 1.0 is standard text-to-audio; <1.0 is a creative effect (model sees pure noise but with t=σmax timestep — mildly OOD); >1.0 is 'guidance overshoot' (more diverse, increasingly weird). --cfg CFG Classifier-Free Guidance scale. 1.0 = off (single forward pass, the rf_denoiser default). >1.0 pushes toward the prompt (classic CFG). [0, 1) pulls toward the unconditional / negative branch (less prompt- aligned, more diverse). <0 actively pushes AWAY from the prompt. Any value ≠ 1.0 costs ~2× per step (batched cond + uncond forward). --apg APG Adaptive Projected Guidance scale [0..1], only matters when --cfg ≠ 1.0. 1.0 = full APG (project the cond−uncond difference orthogonal to cond_denoised; prevents over-saturation at high CFG). 0.0 = vanilla CFG (use the full difference). Intermediate values blend the two. rf_denoiser default is 1.0. --free-models, --no-free-models Free each model after its last use (T5Gemma after step 2, encoder after step 3a, DiT after step 3b) to minimize peak RAM. Lowers decode-stage peak by 3-5 GB on sa3-medium since the DiT (2.7 GB FP16) doesn't sit idle while the decoder runs. Use --no-free-models to keep everything resident — only useful if you plan to call models multiple times in one run (currently the script doesn't). Default: on. --out OUT Output WAV path. Relative paths land in the project's output/ directory (auto-created); absolute paths are used as-is. Always written as 16-bit PCM stereo at 44.1 kHz, trimmed to exactly --seconds. If omitted, auto-named from the prompt and seed. --play After writing the WAV, play it through the default output device via the macOS `afplay` binary. Blocking — the script exits when playback finishes. Ctrl-C stops playback and exits the script (SIGINT is delivered to both processes). modes text-to-audio --prompt P audio-to-audio --prompt P --init-audio IN.wav [--init-noise-level σ] inpainting --prompt P --init-audio IN.wav --inpaint-range START,END negative CFG --prompt P --cfg N --negative-prompt P_NEG run `sa3_mlx.py --help` for per-flag details.