SFX-ENV-20261005-01
Classic music for towns and the GM room; area-specific environment beds and landmark emitters outside towns
Evidence
From logs/evidence/SFX-ENV-20261005-01/. Click a thumbnail for the full image.
Images (2)
Audio (2)
Logs and data (7)
README.md · 2 kB · open
# SFX-ENV-20261005-01 — evidence Task: classic music for towns and the GM room, area-specific environment beds and landmark emitters outside towns (ADR-026). | File | What it is | |---|---| | `audio-check.json` | Objective checks on the 24 delivered environment loops: duration, peak, integrated loudness, loop-seam jump, edge dip, level swing, spectral centroid and flatness. Machine measurements only; they cannot say whether a take sounds good. | | `silent-tail-measurement.json` | The loop defect found on the way: the model ends every clip with 0–1.35 s of digital silence, which the old loop builder looped. Before/after envelope at the seam of `amb-amber-gate`. | | `music-audio-check.json`, `music-structure-check.json`, `music-seam.json` | The same checks on the six delivered music takes (made with the Small Music model): loudness, seam, key match, pulse; plus how softly each take opens and closes. Every take hits its requested tempo (comb fit: 92.0 BPM for Amber Gate, 100.0 for the GM room), so a 60 s loop is a whole number of beats. | | `sfx-model-music-test.json` | Harmonic and rhythmic structure measured on the two music prompts run through the **installed SFX model** (not the music model), next to a white-noise control and a synthetic 92 BPM chord loop. | | `sfx-model-town-theme-test.mp3`, `sfx-model-gm-theme-test.mp3` | The two 30 s clips behind that measurement, so the owner can listen and judge whether the SFX model is good enough for music. Not part of the game audio (`audio/source/`). | Reading `sfx-model-music-test.json`: `keyR` is the best key-profile correlation (noise 0.37, a tonal piece 0.9), `diatonic` the share of pitched energy inside one 7-note scale (noise 0.61), `pulse` the strongest 50–200 BPM regularity of the onset envelope (noise 0.05, the 92 BPM reference 0.78 at 92.3 BPM).
audio-check.json · 9 kB · open
[
{
"file": "audio/source/ambience/amb-amber-gate.mp3",
"seconds": 20.0,
"sampleRate": 44100,
"channels": 2,
"peakDbfs": -3.17,
"rmsDbfs": -29.65,
"lufsIntegrated": -23.3,
"loudnessRangeLu": 4.4,
"seamJump": 0.01327,
"seamLevelStepDb": 0.78,
"edgeDipDb": 0.23,
"levelSwingDb": 3.23,
"spectralCentroidHz": 2970,
"spectralFlatness": 0.071
},
{
"file": "audio/source/ambience/amb-dewgrass-meadow.mp3",
"seconds": 20.0,
"sampleRate": 44100,
"channels": 2,
"peakDbfs": -15.07,
"rmsDbfs": -38.62,
"lufsIntegrated": -30.1,
"loudnessRangeLu": 7.8,
"seamJump": 0.0008,
"seamLevelStepDb": -2.09,
"edgeDipDb": -0.58,
"levelSwingDb": 2.99,
"spectralCentroidHz": 5321,
"spectralFlatness": 0.1022
},
{
"file": "audio/source/ambience/amb-old-orchard.mp3",
"seconds": 20.0,
"sampleRate": 44100,
"channels": 2,
"peakDbfs": -3.03,
"rmsDbfs": -28.74,
…music-audio-check.json · 2 kB · open
[
{
"file": "audio/source/music/mus-amber-gate.mp3",
"seconds": 60.0,
"sampleRate": 44100,
"channels": 2,
"peakDbfs": -3.22,
"rmsDbfs": -19.24,
"lufsIntegrated": -16.7,
"loudnessRangeLu": 3.5,
"seamJump": 0.017,
"seamLevelStepDb": -1.76,
"edgeDipDb": -0.98,
"levelSwingDb": 3.37,
"spectralCentroidHz": 1850,
"spectralFlatness": 0.0275
},
{
"file": "audio/source/music/mus-amber-gate_2.mp3",
"seconds": 60.0,
"sampleRate": 44100,
"channels": 2,
"peakDbfs": -3.11,
"rmsDbfs": -18.13,
"lufsIntegrated": -15.5,
"loudnessRangeLu": 2.4,
"seamJump": 0.06911,
"seamLevelStepDb": 8.36,
"edgeDipDb": 5.94,
"levelSwingDb": 2.72,
"spectralCentroidHz": 1746,
"spectralFlatness": 0.0263
},
{
"file": "audio/source/music/mus-amber-gate_3.mp3",
"seconds": 60.0,
"sampleRate": 44100,
"channels": 2,
"peakDbfs": -3.13,
"rmsDbfs": -18.95,
…music-seam.json · 1 kB · open
{
"note": "mean 100 ms RMS over the first and last 3 s versus the median 100 ms level; a large value means the loop 'breathes' (softer passage) at its loop point",
"takes": [
{
"take": "mus-amber-gate",
"medianDbfs": -20.3,
"first3sDbfs": -21.9,
"last3sDbfs": -19.1,
"seamRegionBelowMedianDb": 1.6
},
{
"take": "mus-amber-gate_2",
"medianDbfs": -18.8,
"first3sDbfs": -20.8,
"last3sDbfs": -18.2,
"seamRegionBelowMedianDb": 2.0
},
{
"take": "mus-amber-gate_3",
"medianDbfs": -19.9,
"first3sDbfs": -21.9,
"last3sDbfs": -20.0,
"seamRegionBelowMedianDb": 2.0
},
{
"take": "mus-gm-proving-ground",
"medianDbfs": -22.0,
"first3sDbfs": -29.7,
"last3sDbfs": -26.6,
"seamRegionBelowMedianDb": 7.6
},
{
"take": "mus-gm-proving-ground_2",
"medianDbfs": -22.5,
"first3sDbfs": -25.2,
"last3sDbfs": -24.0,
"seamRegionBelowMedianDb": 2.8
},
{
"take": "mus-gm-proving-ground_3",
…music-structure-check.json · 972 B · open
[
{
"name": "control: white noise",
"keyR": 0.372,
"diatonic": 0.611,
"pulse": 0.054,
"pulseBpm": 95.7
},
{
"name": "control: synthetic C-major pluck loop at 92 BPM",
"keyR": 0.931,
"diatonic": 0.906,
"pulse": 0.784,
"pulseBpm": 92.3
},
{
"name": "mus-amber-gate.mp3",
"keyR": 0.841,
"diatonic": 0.866,
"pulse": 0.729,
"pulseBpm": 61.5
},
{
"name": "mus-amber-gate_2.mp3",
"keyR": 0.936,
"diatonic": 0.918,
"pulse": 0.598,
"pulseBpm": 61.5
},
{
"name": "mus-amber-gate_3.mp3",
"keyR": 0.713,
"diatonic": 0.792,
"pulse": 0.6,
"pulseBpm": 184.6
},
{
"name": "mus-gm-proving-ground.mp3",
"keyR": 0.664,
"diatonic": 0.934,
…sfx-model-music-test.json · 871 B · open
[
{
"name": "control: white noise",
"keyR": 0.372,
"diatonic": 0.611,
"pulse": 0.054,
"pulseBpm": 95.7
},
{
"name": "control: synthetic C-major pluck loop at 92 BPM",
"keyR": 0.931,
"diatonic": 0.906,
"pulse": 0.784,
"pulseBpm": 92.3
},
{
"name": "sfx-model-mus-amber-gate.mp3",
"keyR": 0.814,
"diatonic": 0.911,
"pulse": 0.104,
"pulseBpm": 136.0
},
{
"name": "sfx-model-mus-gm-proving-ground.mp3",
"keyR": 0.812,
"diatonic": 0.735,
"pulse": 0.275,
"pulseBpm": 198.8
},
{
"name": "amb-cinder-kilns.mp3",
"keyR": 0.492,
"diatonic": 0.706,
"pulse": 0.204,
"pulseBpm": 101.3
},
{
"name": "amb-jade-seal-ruins.mp3",
"keyR": 0.676,
"diatonic": 0.98,
…silent-tail-measurement.json · 1 kB · open
{
"note": "Raw sm-sfx output (cfg 3, negative prompt on) measured in 10 ms blocks; 'content ends' = last block above -70 dBFS. The model ends every clip with digital silence (about -95 dBFS) whose length changes with the requested duration, so the old loop builder (which used the last 0.25 s as its overlap) looped a hole. 2026-10-05.",
"requestedSeconds_contentEnds_silentTailSeconds": [
[6.25, 6.25, 0.0],
[10.25, 10.03, 0.22],
[15.25, 13.9, 1.35],
[20.25, 19.5, 0.75],
[21.25, 20.08, 1.17],
[22.25, 21.36, 0.89],
[30.0, 29.9, 0.1],
[45.25, 44.87, 0.38],
[63.0, 62.54, 0.46]
],
"amberGateBedSeam": {
"note": "100 ms RMS in dBFS around the loop point, amb-amber-gate.mp3",
"before": {
"first1s": [-43.0, -34.0, -29.4, -31.3, -29.4, -28.6, -27.5, -27.9, -29.3, -27.6],
"last1s": [-29.3, -30.9, -30.0, -29.8, -32.3, -77.8, -77.1, -77.3, -77.3, -77.3]
},
"after": {
"first1s": [-32.7, -36.2, -33.6, -33.9, -32.5, -35.0, -37.1, -37.1, -28.5, -30.5],
"last1s": [-31.2, -30.6, -30.0, -30.4, -30.4, -27.2, -27.7, -27.8, -28.8, -32.5]
}
},
"fix": "tools/audio: trim the silent tail before building a loop, request +3 s, refuse a source that is still too short, and reject any finished loop whose first or last 100 ms sits more than 30 dB below the median 100 ms level (edge_dip_db)."
}
Tests
$ /Users/king/AI/stable-audio-3/optimized/mlx/.venv/bin/python -m unittest discover -s tools/audio -p test_*.py$ node --test packages/game-data/tests/feel-quality.test.mjs$ /Users/king/AI/stable-audio-3/optimized/mlx/.venv/bin/python -m unittest discover -s tools/audio -p test_*.py$ node --test packages/game-data/tests/feel-quality.test.mjs$ /Users/king/AI/stable-audio-3/optimized/mlx/.venv/bin/python -m unittest discover -s tools/audio -p test_*.py$ /Users/king/AI/stable-audio-3/optimized/mlx/.venv/bin/python -m unittest discover -s tools/audio -p test_*.py$ pnpm --filter @zoen/wiki typecheck$ pnpm data:build$ pnpm --filter @zoen/wiki typecheck$ /Users/king/AI/stable-audio-3/optimized/mlx/.venv/bin/python -m unittest discover -s tools/audio -p test_*.py$ pnpm --filter @zoen/wiki typecheck$ pnpm --filter @zoen/wiki exec playwright test tests/art.spec.ts -g audio generation$ pnpm --filter @zoen/wiki exec playwright test tests/art.spec.ts -g audio generation$ pnpm --filter @zoen/wiki exec playwright test tests/windows.spec.ts tests/round8.spec.ts -g mapNotes
2026-10-04 18:54 UTC · Finding: the model ends clips with 0–1.35 s of digital silence (varies with the requested length), and sfxgen's loop builder took its body and overlap from inside it, so every ambience loop (including the 12 first-pass ones) had a hole at the seam. Fixing in postprocess/sfxgen (trim tail, ask for +3 s, reject silent loop edges) and regenerating the 24 loops.
2026-10-04 19:33 UTC · Owner request (2026-10-05): music for towns and the GM arena (classic), environment effects outside towns, new prompts in the audio requests, generate as many as needed. Recorded as ADR-026. Agent's reading of 'classic' = classic orchestral fantasy (oud + frame drum kept for Amber Gate); owner can veto. Not decided: boss-fight music, Sand Arena colosseum, day/night variants.
2026-10-04 19:33 UTC · Delivered: 12 area beds rewritten per area (old generic first-pass beds archived in audio/superseded/ambience/2026-10-05-generic-template), 12 landmark spot emitters, 2 music themes (mus-amber-gate 92 BPM, mus-gm-proving-ground 100 BPM; 60 s loops, 3 takes each). Catalog 292 -> 306 prompts, 306/306 delivered. Negative prompts carry 'no music/voices' for beds and emitters; generator stops the build if a zone lacks a written prompt.
2026-10-04 19:33 UTC · Defect found and fixed: the model ends every clip with 0-1.35 s of digital silence (length varies with the requested duration); sfxgen's loop builder looped it, so every ambience loop (incl. the 12 first-pass ones) had a silent hole at the seam. Fix in tools/audio: trim tail before building, request +3 s, refuse a too-short source, reject loops whose first/last 100 ms is >30 dB below the median (edge_dip_db). One regenerated loop was rejected by the new check (37 dB hole) and passed on the seeded retry. Evidence: logs/evidence/SFX-ENV-20261005-01/silent-tail-measurement.json.
2026-10-04 19:33 UTC · Owner asked whether the Small Music model differs from the installed SFX model: yes, separate weights (Stability: small-music 'music-only', small-sfx 'sound-effects-only'). Experiment with the SFX model on both music prompts: key-consistent but no beat (keyR 0.81, pulse 0.10/0.28 vs 92 BPM reference 0.78; noise 0.37/0.05); clips in evidence. Owner then approved the download in chat (dit_sm-music_f16.npz, 919 MB, huggingface.co/stabilityai/stable-audio-3-optimized, anonymous, same licence); installed with ./install.sh -y --python 3.11 --download sm-music, file loads (441 arrays).
2026-10-04 19:33 UTC · Measured, not heard (I cannot listen): 24 loops 20.0 s, peak <= -3 dBFS, edge dip <= 22 dB; 6 music takes 60.0 s, -15.5..-17.5 LUFS, tempo fit exactly 92.0 / 100.0 BPM (60 s = whole beats), GM take 1 opens/closes ~8 dB softer than its median, other five loop level-cleanly. Owner must audition and pick one take per theme.
2026-10-04 19:33 UTC · Also changed: wiki Audio page keeps each delivered sound's prompt/negative prompt/model in a collapsible card section (prompts were only visible on queued rows, so delivered sounds hid them); Maps page shows 'no music, environment sound only' for outdoor areas; e2e audio test no longer assumes a non-empty queue (stale at baseline: queue was already empty) and writes its screenshot to this task's evidence folder; docs-lint exempts apps/wiki/lib/icon-fit.json (pre-existing lint failure); stack.json audio entry and OPEN-3 updated; LICENSES row for Small Music added.
2026-10-04 19:43 UTC · Verification summary: data:test 92/92, sfxgen unit tests 24/24, docs-lint 0 problems, wiki typecheck clean, Playwright audio page (desktop+mobile) and Maps tests pass; browser check of Audio and World pages incl. 360 px (no overflow). First audio e2e run failed on its 120 s timeout (330 files fetched one at a time) and was fixed by batching. NOT run: the full Playwright suite, the whole-site 360 px overflow scan, any listening test (agent cannot hear audio).


