Game guide · source of truth
AI process

Testing & evidence

The test pyramid for Zoen, exact commands, performance and feel tests, video capture, and what counts as proof at each gate.

Zoen is tested at seven levels, from the cheapest to the most expensive. A level never stands in for a higher one: a green unit test does not prove a 1,000-player soak, and a headless bot does not prove a phone renders 60 players.

Test pyramid — cheapest at the bottom, never substitutes upward
7. Human review (clips)
6. Scale / soak (1,000 bots)
5. Performance (frame, net, tick)
4. Browser e2e (Playwright)
3. Deterministic replay
2. Unit / system
1. Data contracts
LevelWhat it provesToolingRuns
1. Data contractsgame rules are consistent (8 picks, no overlapping Atlas nodes, quests valid…)node:test in packages/game-data/testsevery task
2. Unit / systemone function or system is correct (cast phases, AOI sets, damage)Vitest (TS), cargo test (Rust)every task
3. Deterministic replaysame inputs → same server outcome, any frame raterecorded input logs + zoen-sim replayevery combat/server task
4. Browser e2epages/game load, no console errors, interactions workPlaywright with system Chromeevery client/wiki task
5. Performanceframe time, memory, draw calls, bandwidth, tick time inside budgetPlaywright perf harness, criterion, bot loadhot-path tasks + milestone gates
6. Scale / soak1,000 bots per channel for 10–60 min; offline actors for 4 hRust bot swarm + metricsM1, M6, release
7. Human reviewlooks and feels right5–10 s clips, owner sign-offvisual/feel gates

◆Commands

pnpm data:test                                  # level 1: dots and failures; the full report is logs/test-output/data-test-latest.log
pnpm --filter @zoen/client test                 # level 2–4 (client, later)
cargo test --workspace                          # level 2–3 (server, later)
pnpm --filter @zoen/wiki test                   # wiki e2e incl. Atlas smoothness tests (the full suite, at gates)
node tools/verify.mjs <ID> --changed            # per task: the standard checks, then only the wiki specs that changed
node tools/task-log.mjs test <ID> "<name>" -- <any command above>   # always run through the logger

◆Performance tests (how smoothness is measured)

  • Collect frame timestamps with requestAnimationFrame inside the page and report p50/p95/p99, plus the count of frames over 33 ms.
  • Watch for long tasks over 50 ms with PerformanceObserver('longtask').
  • Run on desktop as-is, then with 4× CPU throttling (Chrome DevTools Protocol Emulation.setCPUThrottlingRate) to approximate a mid phone, then on real devices at milestone gates.
  • Drag/pan tests replay a scripted 120-step pointer path. The Atlas tests in this wiki are the template (Ledger shows the results).
  • Game performance tasks name one of the five proposed perf scenes (S1 chain spam, S2 town, S3 world boss, S4 dash traversal, S5 first cast), so runs compare like for like.

◆Feel tests (making "AAA" testable)

CheckHow
VFX/SFX synclog server hit events and client presentation timestamps; assert |Δ| ≤ 16.7 ms
Hitstop doesn't change outcomerun the same replay with feel on/off; compare damage logs byte-for-byte
Reaction choicetable test: impact × weight × poise → reaction (every cell)
Telegraph never culledrender Low with 100 dummies + boss; assert telegraph mesh visible every frame
Input bufferpress Q 120 ms before recovery ends; assert it fires on the first legal frame

◆Video evidence (5–10 s clips)

For battle work, execute C1–C18 control self-tests, the multi-angle animation review and the VFX/environment checks in the guide. Sweep phase boundaries; cover keyboard/mouse, gamepad, touch, both cameras, collision and network faults. Log sampled input → transition → visible motion separately from hardware latency, with command/cast/hit IDs and authority timelines. Agent visual inspection precedes owner review; save pass/fail/not-applicable reasons with frame/time references. A unit test or still capture cannot establish natural motion. Required engine tests remain unverified until a runtime exists.

  • Playwright: browser.newContext({ recordVideo: { dir, size } }) writes WebM. Trim to the action with timestamps in the test.
  • In-page capture: canvas.captureStream(60) + MediaRecorder('video/webm;codecs=vp9') gives frame-accurate gameplay clips (used by the wiki Showcase).
  • Blender previews: tools/blender/*.py render EEVEE clips (Blender bundles FFmpeg, so .mp4 works without a system ffmpeg).
  • Store clips in logs/evidence/<task>/ and reference them in the task note.

◆What a gate needs

A milestone gate is passed only with all of: every task's tests green in logs/tasks.json, the gate metrics recorded (JSON), the clips for visual items, a reviewer scorecard, and the owner's sign-off. Missing platforms are BLOCKED, never "passed by assumption".

Source: zoen/docs/process/TESTING.md · 765 words · edit the Markdown, not this page.