Testing & evidence
The test pyramid for Zoen, exact commands, performance and feel tests, video capture, and what counts as proof at each gate.
Zoen is tested at seven levels, from the cheapest to the most expensive. A level never stands in for a higher one: a green unit test does not prove a 1,000-player soak, and a headless bot does not prove a phone renders 60 players.
| Level | What it proves | Tooling | Runs |
|---|---|---|---|
| 1. Data contracts | game rules are consistent (8 picks, no overlapping Atlas nodes, quests valid…) | node:test in packages/game-data/tests | every task |
| 2. Unit / system | one function or system is correct (cast phases, AOI sets, damage) | Vitest (TS), cargo test (Rust) | every task |
| 3. Deterministic replay | same inputs → same server outcome, any frame rate | recorded input logs + zoen-sim replay | every combat/server task |
| 4. Browser e2e | pages/game load, no console errors, interactions work | Playwright with system Chrome | every client/wiki task |
| 5. Performance | frame time, memory, draw calls, bandwidth, tick time inside budget | Playwright perf harness, criterion, bot load | hot-path tasks + milestone gates |
| 6. Scale / soak | 1,000 bots per channel for 10–60 min; offline actors for 4 h | Rust bot swarm + metrics | M1, M6, release |
| 7. Human review | looks and feels right | 5–10 s clips, owner sign-off | visual/feel gates |
◆Commands
pnpm data:test # level 1: dots and failures; the full report is logs/test-output/data-test-latest.log
pnpm --filter @zoen/client test # level 2–4 (client, later)
cargo test --workspace # level 2–3 (server, later)
pnpm --filter @zoen/wiki test # wiki e2e incl. Atlas smoothness tests (the full suite, at gates)
node tools/verify.mjs <ID> --changed # per task: the standard checks, then only the wiki specs that changed
node tools/task-log.mjs test <ID> "<name>" -- <any command above> # always run through the logger
◆Performance tests (how smoothness is measured)
- Collect frame timestamps with
requestAnimationFrameinside the page and report p50/p95/p99, plus the count of frames over 33 ms. - Watch for long tasks over 50 ms with
PerformanceObserver('longtask'). - Run on desktop as-is, then with 4× CPU throttling (Chrome DevTools Protocol
Emulation.setCPUThrottlingRate) to approximate a mid phone, then on real devices at milestone gates. - Drag/pan tests replay a scripted 120-step pointer path. The Atlas tests in this wiki are the template (Ledger shows the results).
- Game performance tasks name one of the five proposed perf scenes (S1 chain spam, S2 town, S3 world boss, S4 dash traversal, S5 first cast), so runs compare like for like.
◆Feel tests (making "AAA" testable)
| Check | How |
|---|---|
| VFX/SFX sync | log server hit events and client presentation timestamps; assert |Δ| ≤ 16.7 ms |
| Hitstop doesn't change outcome | run the same replay with feel on/off; compare damage logs byte-for-byte |
| Reaction choice | table test: impact × weight × poise → reaction (every cell) |
| Telegraph never culled | render Low with 100 dummies + boss; assert telegraph mesh visible every frame |
| Input buffer | press Q 120 ms before recovery ends; assert it fires on the first legal frame |
◆Video evidence (5–10 s clips)
For battle work, execute C1–C18 control self-tests, the multi-angle animation review and the VFX/environment checks in the guide. Sweep phase boundaries; cover keyboard/mouse, gamepad, touch, both cameras, collision and network faults. Log sampled input → transition → visible motion separately from hardware latency, with command/cast/hit IDs and authority timelines. Agent visual inspection precedes owner review; save pass/fail/not-applicable reasons with frame/time references. A unit test or still capture cannot establish natural motion. Required engine tests remain unverified until a runtime exists.
- Playwright:
browser.newContext({ recordVideo: { dir, size } })writes WebM. Trim to the action with timestamps in the test. - In-page capture:
canvas.captureStream(60)+MediaRecorder('video/webm;codecs=vp9')gives frame-accurate gameplay clips (used by the wiki Showcase). - Blender previews:
tools/blender/*.pyrender EEVEE clips (Blender bundles FFmpeg, so.mp4works without a system ffmpeg). - Store clips in
logs/evidence/<task>/and reference them in the task note.
◆What a gate needs
A milestone gate is passed only with all of: every task's tests green in logs/tasks.json, the gate metrics recorded
(JSON), the clips for visual items, a reviewer scorecard, and the owner's sign-off. Missing platforms are
BLOCKED, never "passed by assumption".
Source: zoen/docs/process/TESTING.md · 765 words · edit the Markdown, not this page.
