Game guide · source of truth
AI process

Development review, round 2

Second review (2026-10-07) of speed, token use and the AAA bar. Measured — the token rules are written but not enforced (the session that wrote them ran at a 425k median context) and the token ledger is losing its history. Proposed — guards that enforce the rules, twelve direction findings (the largest — confirmed hits arrive a round trip late), and 90 new AAA details in four new areas (network feel, input feel, look and rendering, session flow). Decided by the owner on 2026-10-07 (ADR-031).

Review of 2026-10-07 (task REVIEW-20261007-01), decided by the owner the same day (ADR-031): OPEN-15 C, OPEN-16 (a), OPEN-17 yes, OPEN-18 all, OPEN-19 later; the agents make the action AAA and the owner gives feedback later. Section 1 is measured; estimates elsewhere are still estimates. Round 1 is Speed and token efficiency and the AAA detail catalog (ADR-028 to ADR-030); this page builds on them and does not repeat them.

◆Summary

  • The token rules are words, not guards. The session that wrote them ran at a median context of 425k tokens per turn, and 89 of its 95 turns were above the 150k limit it set. Nothing warned it. Section 2 turns the rules into checks.
  • The token ledger is losing its history. It is rebuilt only from transcripts still on disk. Most of them are gone, so the working copy fell from 4,179 messages and 1.97 billion tokens (committed) to 324 messages and 122 million.
  • Confirmed hits arrive a round trip late. Under the accepted layer rule (HIT-04) the impact sound and flash, and on a strict reading the attacker freeze, land 80–130 ms after the contact frame at an 80 ms round trip. Measure it and decide with both versions side by side (D1).
  • Judge feel at real latency. A first skill approved at 0 ms in the Combat Lab can feel soft online. Review at a simulated 80 ms (NET-01).
  • 90 new AAA details, six of them core, in four new areas: network feel, input feel, look and rendering, and session flow. A proposed fourth stage, alpha, keeps the G3 battle slice small.
  • Two conflicts fixed: UI-05 now uses the HUD's 48 px touch targets (hud.json), and the accepted core count is 38, not 37 (WLD-08 moved to core on 2026-10-06).

◆1. Measured since round 1

MeasuredValueSource
Round-1 session (the one that wrote the token rules)95 turns; median context 425k, 90th percentile 573k, largest 606k; 89 of 95 turns over 150ktranscript 55692815…
Its tool calls83 shell, 24 file reads, 13 web fetches, 8 web searches, 10 file writessame
Art queue session219 turns at a median context of 404k; 63 browser actions and 50 page scriptstranscript 25ebee33…
Token ledger, committed (2026-10-05)4,179 messages · 1,974 M tokensgit show HEAD:logs/token-ledger.json
Token ledger, working copy (2026-10-06)324 messages · 122 M tokens: the older transcripts are no longer on disk, and the rebuild dropped themlogs/token-ledger.json
Logged test runs754 runs, 167 failed (22%)logs/tasks.jsonl
Slow loopsfull wiki browser suite about 8 min per run (4 runs); one site suite 29 minsame

What follows from it. Cost is still turns × context. Round 1 made the rules right; it did not make them automatic, and a long session never notices that it has become expensive. The ledger, the tool that should show the problem, now hides most of it. The fixes below are mechanical so that no agent has to remember them.

◆2. Make the rules enforce themselves

#ChangeEffectStatus
E1The ledger keeps its history. Merge each rebuild with the previous ledger, never drop a session whose transcript is gone, and restore 2026-10-03 to 2026-10-05 from gitTotals stop shrinking; task rows survive transcript clean-upDone (LEDGER-20261007-01): merge on every rebuild, history restored from 3465c40, sub-agent transcripts counted (how)
E2Read guard. A Claude Code PreToolUse hook refuses whole reads of skills.json, audio-requests.json, art-requests*.json, art-library.json, atlas.json and logs/tasks.json (later also Unity scene and prefab files over 200 KB) and answers with the ctx.mjs command to useA rule becomes a guard. One blocked read saves 30k–330k tokens on every later turn of that sessionDone (ENFORCE-20261007-01): tools/guard/read-guard.mjs, a PreToolUse hook (tools)
E3Context meter. After each tool call a hook reads the session's last usage; at 150k it tells the agent once, at 250k it asks it to finish and hand offThe agent learns its cost while it can still act on itDone (ENFORCE-20261007-01): tools/guard/context-meter.mjs, a PostToolUse hook
E4Guide slices. ctx.mjs guide <section> prints one section of the development guide (its {#anchor} ids), and each roadmap task lists the sections it needsThe mandatory read falls from about 8.5k tokens to 1–2k for most tasksDone (CTX-GUIDE-20261007-01): ctx.mjs guide and ctx.mjs doc print one {#anchor} section, and 14 M0–M2 tasks name theirs in roadmap.json (median 3.4 KB, about 0.8k tokens, against 35.1 KB, about 8.8k tokens, for the whole guide)
E5Every finish opens the next session. The finishing agent creates the next smallest action as a new desktop-app session with a self-contained prompt; the owner clicks onceA fresh session becomes the easy path instead of a long oneDone (ENFORCE-20261007-01): in the kickoff prompt and the task loop
E6Quiet by default. Data tests print dots and failures, while the full report still goes to logs/test-output/; Playwright runs changed specs per task (--only-changed) and the full suites at gatesSmaller outputs; fewer 8–29 minute runsDone (QUIET-TESTS-20261007-01): tools/quiet-reporter.mjs makes pnpm data:test print 573 bytes instead of 31,409 (136 tests; full spec report in logs/test-output/data-test-latest.log), and verify.mjs <ID> --changed runs only the changed Playwright specs (tools)
E7A model per lane. Agent definitions in zoen/.claude/agents/: an explorer on the small model, a queue runner on the middle model, a read-only reviewer on the strongestADR-029 rule 7 happens by defaultDone (ENFORCE-20261007-01): explorer, queue-runner and reviewer
E8State a rule once. docs-lint reports any sentence repeated in three or more docsThe movement-cancel rule is restated in more than ten files; each copy is read and can driftLater task
E9No known failures. Fix the three failing round-8 inventory tests or mark them fixme with a reasonA run with known failures costs an investigation every timeSeparate task

◆3. Faster game development, round 2

These add to changes A–H of round 1.

#ChangeWhy it is faster
F1Judge at latency (NET-01, NET-02). The Combat Lab defaults to an 80 ms simulated round tripFeel approved at 0 ms would be re-tuned after M6; this tunes it once
F2Live tuning. The Combat Lab reloads feel values (hitstop, shake, cue offsets, gains) from the JSON while it runs; a GM tuning panel exports a JSON patch that the agent applies to canonical dataThe owner tunes twenty values in one sitting instead of twenty agent turns and rebuilds
F3Timeline in the wiki. The accepted combat timeline debugger writes JSON (phases, cues, hits, corrections), and a wiki page draws itAgents read the numbers and the owner reads the picture; a two-frame offset shows without a video
F4Types from data. JSON Schema for the files the runtime reads; Rust and C# types generated from it (for example quicktype, Apache-2.0, after the toolkit gate)No agent writes the data model three times; a schema change is one diff
F5Web-safe lint, Web build at gates. A static check keeps compute shaders, VFX Graph, threads and mixer effects out of shared runtime code; the browser build runs at milestone gates and nightly once a runner existsTasks stop paying the browser tax each time; each gate still proves the browser
F6Owner-performed motion (optional, OPEN-19). The owner films signature moves (as for the Wukong walk); single-camera video-to-motion gives base curves that the agent cleans upSkill-specific motion that Mixamo lacks; owner review stays the bottleneck either way
F7Catalog coverage from the task log. Task notes already name detail ids; a script counts done and in progress per stage for the wikiThe owner sees "core 0/44" without asking an agent

◆4. Direction findings

◆D1 · Confirmed hits arrive a round trip late

The server starts a cast when the input arrives and resolves contact on a 20 Hz tick. Its Hit event comes back one round trip plus up to one tick after the local contact frame. Client prediction does not help the target layer, and backdating is ruled out (cancel contract).

Round tripConfirmed impact after the local contact frameFrames at 30 fpsFrames at 60 fps
60 ms60–110 ms2–34–7
80 ms (proposed review latency)80–130 ms2–45–8
120 ms120–170 ms4–57–10
200 ms200–250 ms6–812–15

Arithmetic only, before snapshot batching and display latency. HIT-04 predicts only the wind-up, swish, trail and cast cue. Read strictly, the attacker's freeze (HIT-01) and the impact sound wait for the server, so online the freeze holds a follow-through pose instead of the contact pose that ANM-06 asks for, and the impact trails the swish peak (SFX-02). NETCODE_1000.md still allows predicted contact sparks; round 1 settled that in favour of HIT-04.

  • A · Keep HIT-04 strict. Simplest and honest about misses; hits feel softer as latency grows.
  • B · Predicted contact. At the contact frame the client's own copy of the hit test decides. That copy is the shared simulation of ADR-029 and spike M0-11, run against the lag-compensated target positions it already draws. On a hit, the attacker freeze, spark and first impact transient play at once. Numbers, health, reactions, status and kill confirm still wait for the server. A rejected prediction shows a spark without a number, and telemetry counts it.
  • C · Measure, then choose (recommended). Build A and B behind a Combat Lab switch, record the delay (NET-02) and let the owner judge both at 80 ms in the G2 review. B is the likely winner if M0-11 succeeds.

◆Other findings

#FindingProposal
D2A first skill judged at 0 ms passes G2 and can still feel soft onlineNET-01 (core): reviews and the G2 clip at a simulated 80 ms round trip
D3Every technique needs a browser fallback, so the browser taxes every taskF5; keep the browser at Low–Mid (existing); re-check WebGPU support in the pinned editor before relying on it
D4One Animator per character cannot carry 1,000 players and 4,000 monsters on LowOPT-20 update rates by distance, on top of the baked crowds of M1-01
D5A per-character hit flash done with MaterialPropertyBlocks takes every flashing character out of the SRP BatcherOPT-19 (core): material instances or instance data. The first skill's hit flash must use it from day one
D6Bloom thresholds and effect palettes depend on the tone mapper and exposure, which nothing fixes yetLOOK-01 (core): one colour pipeline, chosen in the Combat Lab before the first effect review
D7UI-05 gave 44 px touch targets; the HUD contract says 48 pxFixed: UI-05 cites hud.json
D8The ledger drops history (section 1)E1
D9Accepted details already need voices: a hurt vocal (HIT-08), aggro barks (ENM-04), creature voices (SFX-15). No voice source is chosenOPEN-16
D10The first minutes of play, death and returning after a break have no specificationFLOW area and the alpha stage
D11Agents cannot judge motion (ADR-029), so animation review is the owner's bottleneck for 144 skillsThe skill factory (round 1, D), contact sheets, and optionally F6
D1222% of logged test runs failed, and the slowest suites take 8–29 minutesE6 and E9; run the cheap checks before the slow ones (verify.mjs already does)

◆5. AAA detail catalog, round 2

90 details accepted on 2026-10-07 (ADR-031), in the catalog with their added date. They change no server outcome and no accepted number. Where they meet existing rules they cite them (hud.json, NETCODE_1000.md, PERFORMANCE.md, controls.json, feel-quality.json).

AreaNewHighlights
Hit feel4Ticks are not hits (HIT-14), critical read, projectile impacts, finisher signature
Camera3Fit the fight for bosses, the view-switch blend, camera volumes for interiors
Animation6Cues come from the cast clock (ANM-14, core), draw and sheathe without delay, landing weight, no lockstep crowds
Effects5Spawn and despawn, decals that stay on the ground, far effects simplify, one shared texture library
Sound7Borders blend, big sources, no Doppler wobble, focus and interruptions, player efforts, after-the-fight release
World4One wind, NPCs notice you, things you can use, footprints
Enemies2Dead means untargeted, boss presence
Loot1New and better markers
Interface8Device glyphs, gamepad windows, steady numbers, errors that help, cursor, drag and drop
Optimisation7Batching-safe character effects (OPT-19, core), animation by distance, cosmetic-only client physics, texture formats, per-system timings
Comfort6Text size, captions, hold or toggle, high-contrast threats, close-view field of view
Network feel (new)10Judge at latency and measure the confirm delay (NET-01, NET-02, core), others move like locals, elastic buffer, reconnect
Input feel (new)8Stick shaping (INP-01, core), touch stick and aiming, device hot-swap, haptics vocabulary, cursor aim on slopes
Look and rendering (new)12One colour pipeline (LOOK-01, core), material and texel checks, characters stand out, anti-aliasing by tier, no banding
Session flow (new)7Boot to play, the first minute, taught by doing, death and return, arrivals

By stage: 6 core, 44 slice, 27 alpha and 13 polish; core grew from 38 to 44 details.

Why a fourth stage. Without it, every detail that is not core lands in slice and the G3 gate would carry about 119 details. The alpha stage (needed by M8, the first closed playtest) holds online feel, onboarding, full menus and comfort options, so G3 stays a small real fight. Agents load one stage with node tools/ctx.mjs polish alpha.

◆6. Decisions

IDQuestionOptionsRecommendation · default until answeredOwner, 2026-10-07
OPEN-15How should hits feel online (D1)?A strict confirmed layers · B predicted contact · C build both and judge at 80 msC, B expected · default A with NET-02 measuringC
OPEN-16Where do voices come from (D9)?(a) non-verbal efforts and creature voices from the local generator · (b) text-to-speech barks after a licence check · (c) recorded by the owner or volunteers · (d) a licensed pack inside the PHP 5,000 cap(a) now, barks revisited at M8 · default (a)(a)
OPEN-17Enable the enforcement pack: read guard (E2), context meter (E3) and lane agents (E7)?Yes · noYes. Repository files only, removable · default offYes
OPEN-18Accept the round-2 catalog (90 details, 6 of them core) and the alpha stage?All · some · noneAll · default: proposals only; the 38 accepted core details still gate G2All
OPEN-19Try video-to-motion for signature skills (F6)? Needs an account and a tool comparisonYes · later · noLater, after the mannequin (M2-09) · default Mixamo plus agent polishLater; make the action AAA, feedback later

◆Sources

Checked 2026-10-07. Each source supports only the fact beside it.

  • Unity manual, SRP Batcher compatibility: a GameObject that uses MaterialPropertyBlocks is not SRP Batcher compatible (OPT-19).
  • Unity manual, texture compression in Web: one format per build (DXT for desktop browsers, ASTC for mobile); an unsupported format is decompressed in software (OPT-22).
  • Playwright, command line: --only-changed runs the test files changed since a git reference (E6).
  • Node.js, test reporters: the dot reporter, and several reporters with separate destinations (E6).

Source: zoen/docs/process/REVIEW_ROUND_2.md · 2,915 words · edit the Markdown, not this page.