Development review, round 2
Second review (2026-10-07) of speed, token use and the AAA bar. Measured — the token rules are written but not enforced (the session that wrote them ran at a 425k median context) and the token ledger is losing its history. Proposed — guards that enforce the rules, twelve direction findings (the largest — confirmed hits arrive a round trip late), and 90 new AAA details in four new areas (network feel, input feel, look and rendering, session flow). Decided by the owner on 2026-10-07 (ADR-031).
Review of 2026-10-07 (task
REVIEW-20261007-01), decided by the owner the same day (ADR-031): OPEN-15 C, OPEN-16 (a), OPEN-17 yes, OPEN-18 all, OPEN-19 later; the agents make the action AAA and the owner gives feedback later. Section 1 is measured; estimates elsewhere are still estimates. Round 1 is Speed and token efficiency and the AAA detail catalog (ADR-028 to ADR-030); this page builds on them and does not repeat them.
◆Summary
- The token rules are words, not guards. The session that wrote them ran at a median context of 425k tokens per turn, and 89 of its 95 turns were above the 150k limit it set. Nothing warned it. Section 2 turns the rules into checks.
- The token ledger is losing its history. It is rebuilt only from transcripts still on disk. Most of them are gone, so the working copy fell from 4,179 messages and 1.97 billion tokens (committed) to 324 messages and 122 million.
- Confirmed hits arrive a round trip late. Under the accepted layer rule (HIT-04) the impact sound and flash, and on a strict reading the attacker freeze, land 80–130 ms after the contact frame at an 80 ms round trip. Measure it and decide with both versions side by side (D1).
- Judge feel at real latency. A first skill approved at 0 ms in the Combat Lab can feel soft online. Review at a simulated 80 ms (NET-01).
- 90 new AAA details, six of them
core, in four new areas: network feel, input feel, look and rendering, and session flow. A proposed fourth stage,alpha, keeps the G3 battle slice small. - Two conflicts fixed: UI-05 now uses the HUD's 48 px touch targets (hud.json), and the accepted
corecount is 38, not 37 (WLD-08 moved tocoreon 2026-10-06).
◆1. Measured since round 1
| Measured | Value | Source |
|---|---|---|
| Round-1 session (the one that wrote the token rules) | 95 turns; median context 425k, 90th percentile 573k, largest 606k; 89 of 95 turns over 150k | transcript 55692815… |
| Its tool calls | 83 shell, 24 file reads, 13 web fetches, 8 web searches, 10 file writes | same |
| Art queue session | 219 turns at a median context of 404k; 63 browser actions and 50 page scripts | transcript 25ebee33… |
| Token ledger, committed (2026-10-05) | 4,179 messages · 1,974 M tokens | git show HEAD:logs/token-ledger.json |
| Token ledger, working copy (2026-10-06) | 324 messages · 122 M tokens: the older transcripts are no longer on disk, and the rebuild dropped them | logs/token-ledger.json |
| Logged test runs | 754 runs, 167 failed (22%) | logs/tasks.jsonl |
| Slow loops | full wiki browser suite about 8 min per run (4 runs); one site suite 29 min | same |
What follows from it. Cost is still turns × context. Round 1 made the rules right; it did not make them automatic, and a long session never notices that it has become expensive. The ledger, the tool that should show the problem, now hides most of it. The fixes below are mechanical so that no agent has to remember them.
◆2. Make the rules enforce themselves
| # | Change | Effect | Status |
|---|---|---|---|
| E1 | The ledger keeps its history. Merge each rebuild with the previous ledger, never drop a session whose transcript is gone, and restore 2026-10-03 to 2026-10-05 from git | Totals stop shrinking; task rows survive transcript clean-up | Done (LEDGER-20261007-01): merge on every rebuild, history restored from 3465c40, sub-agent transcripts counted (how) |
| E2 | Read guard. A Claude Code PreToolUse hook refuses whole reads of skills.json, audio-requests.json, art-requests*.json, art-library.json, atlas.json and logs/tasks.json (later also Unity scene and prefab files over 200 KB) and answers with the ctx.mjs command to use | A rule becomes a guard. One blocked read saves 30k–330k tokens on every later turn of that session | Done (ENFORCE-20261007-01): tools/guard/read-guard.mjs, a PreToolUse hook (tools) |
| E3 | Context meter. After each tool call a hook reads the session's last usage; at 150k it tells the agent once, at 250k it asks it to finish and hand off | The agent learns its cost while it can still act on it | Done (ENFORCE-20261007-01): tools/guard/context-meter.mjs, a PostToolUse hook |
| E4 | Guide slices. ctx.mjs guide <section> prints one section of the development guide (its {#anchor} ids), and each roadmap task lists the sections it needs | The mandatory read falls from about 8.5k tokens to 1–2k for most tasks | Done (CTX-GUIDE-20261007-01): ctx.mjs guide and ctx.mjs doc print one {#anchor} section, and 14 M0–M2 tasks name theirs in roadmap.json (median 3.4 KB, about 0.8k tokens, against 35.1 KB, about 8.8k tokens, for the whole guide) |
| E5 | Every finish opens the next session. The finishing agent creates the next smallest action as a new desktop-app session with a self-contained prompt; the owner clicks once | A fresh session becomes the easy path instead of a long one | Done (ENFORCE-20261007-01): in the kickoff prompt and the task loop |
| E6 | Quiet by default. Data tests print dots and failures, while the full report still goes to logs/test-output/; Playwright runs changed specs per task (--only-changed) and the full suites at gates | Smaller outputs; fewer 8–29 minute runs | Done (QUIET-TESTS-20261007-01): tools/quiet-reporter.mjs makes pnpm data:test print 573 bytes instead of 31,409 (136 tests; full spec report in logs/test-output/data-test-latest.log), and verify.mjs <ID> --changed runs only the changed Playwright specs (tools) |
| E7 | A model per lane. Agent definitions in zoen/.claude/agents/: an explorer on the small model, a queue runner on the middle model, a read-only reviewer on the strongest | ADR-029 rule 7 happens by default | Done (ENFORCE-20261007-01): explorer, queue-runner and reviewer |
| E8 | State a rule once. docs-lint reports any sentence repeated in three or more docs | The movement-cancel rule is restated in more than ten files; each copy is read and can drift | Later task |
| E9 | No known failures. Fix the three failing round-8 inventory tests or mark them fixme with a reason | A run with known failures costs an investigation every time | Separate task |
◆3. Faster game development, round 2
These add to changes A–H of round 1.
| # | Change | Why it is faster |
|---|---|---|
| F1 | Judge at latency (NET-01, NET-02). The Combat Lab defaults to an 80 ms simulated round trip | Feel approved at 0 ms would be re-tuned after M6; this tunes it once |
| F2 | Live tuning. The Combat Lab reloads feel values (hitstop, shake, cue offsets, gains) from the JSON while it runs; a GM tuning panel exports a JSON patch that the agent applies to canonical data | The owner tunes twenty values in one sitting instead of twenty agent turns and rebuilds |
| F3 | Timeline in the wiki. The accepted combat timeline debugger writes JSON (phases, cues, hits, corrections), and a wiki page draws it | Agents read the numbers and the owner reads the picture; a two-frame offset shows without a video |
| F4 | Types from data. JSON Schema for the files the runtime reads; Rust and C# types generated from it (for example quicktype, Apache-2.0, after the toolkit gate) | No agent writes the data model three times; a schema change is one diff |
| F5 | Web-safe lint, Web build at gates. A static check keeps compute shaders, VFX Graph, threads and mixer effects out of shared runtime code; the browser build runs at milestone gates and nightly once a runner exists | Tasks stop paying the browser tax each time; each gate still proves the browser |
| F6 | Owner-performed motion (optional, OPEN-19). The owner films signature moves (as for the Wukong walk); single-camera video-to-motion gives base curves that the agent cleans up | Skill-specific motion that Mixamo lacks; owner review stays the bottleneck either way |
| F7 | Catalog coverage from the task log. Task notes already name detail ids; a script counts done and in progress per stage for the wiki | The owner sees "core 0/44" without asking an agent |
◆4. Direction findings
◆D1 · Confirmed hits arrive a round trip late
The server starts a cast when the input arrives and resolves contact on a 20 Hz tick. Its Hit event comes back one round trip plus up to one tick after the local contact frame. Client prediction does not help the target layer, and backdating is ruled out (cancel contract).
| Round trip | Confirmed impact after the local contact frame | Frames at 30 fps | Frames at 60 fps |
|---|---|---|---|
| 60 ms | 60–110 ms | 2–3 | 4–7 |
| 80 ms (proposed review latency) | 80–130 ms | 2–4 | 5–8 |
| 120 ms | 120–170 ms | 4–5 | 7–10 |
| 200 ms | 200–250 ms | 6–8 | 12–15 |
Arithmetic only, before snapshot batching and display latency. HIT-04 predicts only the wind-up, swish, trail and cast cue. Read strictly, the attacker's freeze (HIT-01) and the impact sound wait for the server, so online the freeze holds a follow-through pose instead of the contact pose that ANM-06 asks for, and the impact trails the swish peak (SFX-02). NETCODE_1000.md still allows predicted contact sparks; round 1 settled that in favour of HIT-04.
- A · Keep HIT-04 strict. Simplest and honest about misses; hits feel softer as latency grows.
- B · Predicted contact. At the contact frame the client's own copy of the hit test decides. That copy is the shared simulation of ADR-029 and spike M0-11, run against the lag-compensated target positions it already draws. On a hit, the attacker freeze, spark and first impact transient play at once. Numbers, health, reactions, status and kill confirm still wait for the server. A rejected prediction shows a spark without a number, and telemetry counts it.
- C · Measure, then choose (recommended). Build A and B behind a Combat Lab switch, record the delay (NET-02) and let the owner judge both at 80 ms in the G2 review. B is the likely winner if M0-11 succeeds.
◆Other findings
| # | Finding | Proposal |
|---|---|---|
| D2 | A first skill judged at 0 ms passes G2 and can still feel soft online | NET-01 (core): reviews and the G2 clip at a simulated 80 ms round trip |
| D3 | Every technique needs a browser fallback, so the browser taxes every task | F5; keep the browser at Low–Mid (existing); re-check WebGPU support in the pinned editor before relying on it |
| D4 | One Animator per character cannot carry 1,000 players and 4,000 monsters on Low | OPT-20 update rates by distance, on top of the baked crowds of M1-01 |
| D5 | A per-character hit flash done with MaterialPropertyBlocks takes every flashing character out of the SRP Batcher | OPT-19 (core): material instances or instance data. The first skill's hit flash must use it from day one |
| D6 | Bloom thresholds and effect palettes depend on the tone mapper and exposure, which nothing fixes yet | LOOK-01 (core): one colour pipeline, chosen in the Combat Lab before the first effect review |
| D7 | UI-05 gave 44 px touch targets; the HUD contract says 48 px | Fixed: UI-05 cites hud.json |
| D8 | The ledger drops history (section 1) | E1 |
| D9 | Accepted details already need voices: a hurt vocal (HIT-08), aggro barks (ENM-04), creature voices (SFX-15). No voice source is chosen | OPEN-16 |
| D10 | The first minutes of play, death and returning after a break have no specification | FLOW area and the alpha stage |
| D11 | Agents cannot judge motion (ADR-029), so animation review is the owner's bottleneck for 144 skills | The skill factory (round 1, D), contact sheets, and optionally F6 |
| D12 | 22% of logged test runs failed, and the slowest suites take 8–29 minutes | E6 and E9; run the cheap checks before the slow ones (verify.mjs already does) |
◆5. AAA detail catalog, round 2
90 details accepted on 2026-10-07 (ADR-031), in the catalog with their added date. They change no server outcome and no accepted number. Where they meet existing rules they cite them (hud.json, NETCODE_1000.md, PERFORMANCE.md, controls.json, feel-quality.json).
| Area | New | Highlights |
|---|---|---|
| Hit feel | 4 | Ticks are not hits (HIT-14), critical read, projectile impacts, finisher signature |
| Camera | 3 | Fit the fight for bosses, the view-switch blend, camera volumes for interiors |
| Animation | 6 | Cues come from the cast clock (ANM-14, core), draw and sheathe without delay, landing weight, no lockstep crowds |
| Effects | 5 | Spawn and despawn, decals that stay on the ground, far effects simplify, one shared texture library |
| Sound | 7 | Borders blend, big sources, no Doppler wobble, focus and interruptions, player efforts, after-the-fight release |
| World | 4 | One wind, NPCs notice you, things you can use, footprints |
| Enemies | 2 | Dead means untargeted, boss presence |
| Loot | 1 | New and better markers |
| Interface | 8 | Device glyphs, gamepad windows, steady numbers, errors that help, cursor, drag and drop |
| Optimisation | 7 | Batching-safe character effects (OPT-19, core), animation by distance, cosmetic-only client physics, texture formats, per-system timings |
| Comfort | 6 | Text size, captions, hold or toggle, high-contrast threats, close-view field of view |
| Network feel (new) | 10 | Judge at latency and measure the confirm delay (NET-01, NET-02, core), others move like locals, elastic buffer, reconnect |
| Input feel (new) | 8 | Stick shaping (INP-01, core), touch stick and aiming, device hot-swap, haptics vocabulary, cursor aim on slopes |
| Look and rendering (new) | 12 | One colour pipeline (LOOK-01, core), material and texel checks, characters stand out, anti-aliasing by tier, no banding |
| Session flow (new) | 7 | Boot to play, the first minute, taught by doing, death and return, arrivals |
By stage: 6 core, 44 slice, 27 alpha and 13 polish; core grew from 38 to 44 details.
Why a fourth stage. Without it, every detail that is not core lands in slice and the G3 gate would carry about
119 details. The alpha stage (needed by M8, the first closed playtest) holds online feel, onboarding, full
menus and comfort options, so G3 stays a small real fight. Agents load one stage with
node tools/ctx.mjs polish alpha.
◆6. Decisions
| ID | Question | Options | Recommendation · default until answered | Owner, 2026-10-07 |
|---|---|---|---|---|
| OPEN-15 | How should hits feel online (D1)? | A strict confirmed layers · B predicted contact · C build both and judge at 80 ms | C, B expected · default A with NET-02 measuring | C |
| OPEN-16 | Where do voices come from (D9)? | (a) non-verbal efforts and creature voices from the local generator · (b) text-to-speech barks after a licence check · (c) recorded by the owner or volunteers · (d) a licensed pack inside the PHP 5,000 cap | (a) now, barks revisited at M8 · default (a) | (a) |
| OPEN-17 | Enable the enforcement pack: read guard (E2), context meter (E3) and lane agents (E7)? | Yes · no | Yes. Repository files only, removable · default off | Yes |
| OPEN-18 | Accept the round-2 catalog (90 details, 6 of them core) and the alpha stage? | All · some · none | All · default: proposals only; the 38 accepted core details still gate G2 | All |
| OPEN-19 | Try video-to-motion for signature skills (F6)? Needs an account and a tool comparison | Yes · later · no | Later, after the mannequin (M2-09) · default Mixamo plus agent polish | Later; make the action AAA, feedback later |
◆Sources
Checked 2026-10-07. Each source supports only the fact beside it.
- Unity manual, SRP Batcher compatibility: a GameObject that uses MaterialPropertyBlocks is not SRP Batcher compatible (OPT-19).
- Unity manual, texture compression in Web: one format per build (DXT for desktop browsers, ASTC for mobile); an unsupported format is decompressed in software (OPT-22).
- Playwright, command line:
--only-changedruns the test files changed since a git reference (E6). - Node.js, test reporters: the
dotreporter, and several reporters with separate destinations (E6).
Source: zoen/docs/process/REVIEW_ROUND_2.md · 2,915 words · edit the Markdown, not this page.
