Speed and token efficiency
Where the tokens went in the first three days (measured from the transcripts), the rules and tools that cut the cost of every later task, and the changes that would make game development itself faster — work that can start before Unity is installed, one simulation instead of two, a skill factory, and numeric checks before visual review. Accepted by the owner on 2026-10-06 (ADR-029).
Review of 2026-10-06, accepted by the owner the same day (ADR-029). Section 1 is measured. Sections 2, 4 and 5 are now rules and planned work; the estimates in them are still estimates. Prices are the list prices in
tools/pricing.json.
◆1. Where the tokens went
Read from the Claude Code transcripts of 2026-10-03 to 2026-10-05, each assistant turn counted once.
| Measured | Value |
|---|---|
| Turns · sessions · time | 4,096 turns in 12 sessions; 55.6 h wall clock, 26.8 h active |
| Cost at list prices | USD 557 |
| Split of that cost | context re-read from cache 69% (USD 384) · output 18% (USD 101) · cache writes 13% (USD 73) |
| Context re-read per turn | median 465k tokens, 90th percentile 829k, largest 966k |
| Share of cost in turns above 150k context | 96% |
| Two longest sessions (28 h and 12.5 h, never cleared) | 78% of all cost |
| Tasks over 60 minutes | 18 of 229; the longest ran 9 h |
| Sum of per-task costs in the Ledger | USD 1,039, which is 1.86× the real total |
Two conclusions follow.
- Cost is turns × context size. Writing is cheap; re-reading a huge conversation on every tool call is not. At the median context, one extra tool call cost about USD 0.09 before it did anything.
- Per-task numbers were double-counted. 156 of 229 tasks carried no session id, so two sessions working at the same time were both counted in each other's tasks. The Ledger's grand total was always right; its task rows were not.
◆2. Token rules
Binding since 2026-10-06; the short form is in CLAUDE.md. Ordered by effect.
- One task, one fresh session. Start a new session for each task card and end it at
finish. The next session starts fromCONTINUE_HERE.mdand the card, not from an old conversation. Keep the average context under 150k tokens. Estimate: the same 4,096 turns at an 80k average would have cost about USD 66 in cache reads instead of 384; with the extra cache writes of fresh starts, the total falls by roughly half. - Load slices, not files. Use
node tools/ctx.mjs(section 3). Never readskills.json(about 85k tokens),audio-requests.json(about 100k) orart-requests.json(about 330k) whole: a file read once is re-read on every later turn of that session. - Batch commands. One script that runs the checks and prints a ten-line summary replaces five tool calls. Full
output already goes to
logs/test-output/; read it only on failure. - Loops are scripts, not turns. Queues (art batches, audio takes, validators over 144 skills) run inside one command. The agent reads the summary, not every item.
- Numbers first, pictures last. A validator's JSON costs a few hundred tokens; a full-size frame costs on the order of a thousand, and an eight-view review of one skill is dozens of frames. Run the numeric checks first, then look at one contact sheet, then only at the frames a check flagged (section 5).
- Wide searches go to a sub-agent. Its context is separate; only its conclusion returns.
- Match the model to the work. Architecture, netcode, feel tuning and reviews need the strongest model.
Routine implementation and data edits do not, and queue supervision needs the least. In
tools/pricing.jsonoutput on the top model is priced up to 10 times the smallest model's, while cached context differs by at most 2.5 times, so rule 1 matters more than this one. - Keep the read-first set small. Before this review an agent loaded about 25k tokens before any task-specific
reading:
CONTINUE_HERE.md(6.4k, mostly round history), the development guide (8k), the index (3.9k) and the standards data (2.9k). Since 2026-10-06CONTINUE_HERE.mdholds status and next actions only, and its history is in the changelog. Still open: the movement-cancel rule is restated in at least ten files; state a rule once in data and link to it.
◆3. Tools that exist now
| Tool | What it does |
|---|---|
node tools/ctx.mjs skill <id> | One skill's contract, frame marks, impact class, weapon family, palette and sound status: about 4 KB instead of 291 KB. Also monster, zone, polish, task, find and sizes. |
node tools/ctx.mjs guide <section> | One section of the development guide by its {#anchor} id, for example cancel-contract: 1 to 3 KB instead of 35 KB (about 8.8k tokens; sizes are bytes ÷ 4). A ## section includes its ### children, several ids print in one call, and guide alone lists every section with its size. The M0 to M2 tasks that need the guide name their sections in a guide array in roadmap.json, and ctx.mjs task <ID> prints the exact command: the 14 cards that do have a median slice of 3.4 KB and a largest of 10.3 KB (M2-10). Read the whole guide only for direction work. node tools/ctx.mjs doc <path>#<anchor> slices any doc under docs/ the same way (a {#id} or the wiki's heading slug). |
node tools/task-log.mjs | Now attributes a task to the running Claude Code session by itself, counts every session that touched the task, and prints turns and average context at finish, with a warning above 150k. |
node tools/verify.mjs <ID> [--changed] | The standard checks of a docs, data or wiki task in one command (data tests, docs lint, wiki typecheck, wiki build), each step logged, one line per step. --changed adds the wiki Playwright run with --only-changed: only the spec files with uncommitted changes (new ones too) and the specs that import a changed file, so a task that touches no spec runs none and passes. --e2e tests/<spec> runs one named spec instead. The full suite (pnpm --filter @zoen/wiki test, 8 to 29 minutes) is for milestone gates. |
Quiet data tests (tools/quiet-reporter.mjs, E6) | pnpm data:test prints one dot per passing test, an X per failure, then each failure with its assertion message (and the stderr of a test file that could not load), then the counts as # pass N and # fail N: 573 bytes instead of 31,409 for 136 tests, most of it pnpm's own header. The full spec report goes to logs/test-output/data-test-latest.log (git-ignored, rewritten every run); read it only on failure. packages/game-data/tests/quiet-reporter.test.mjs proves that a failure stays visible. |
| Data-rendered pages | The AAA detail catalog renders its tables from polish-standards.json. A number lives in one file, so it cannot drift between pages. |
Read guard (tools/guard/read-guard.mjs, ADR-031) | A PreToolUse hook on Read and Bash. It refuses whole reads of skills.json, audio-requests.json, art-requests.json, art-requests.manual.json, art-library.json, atlas.json and logs/tasks.json, and of Unity .unity and .prefab files over 200 KB, and answers with the ctx.mjs or jq slice to use. A whole read is a Read without a limit of at most 500 lines, or cat, less, more, a big head/tail or jq . whose output is not piped into a filter or redirected to a file. Slices pass: grep, jq filters, node and tools/*.mjs. |
Context meter (tools/guard/context-meter.mjs, ADR-031) | A PostToolUse hook. After each tool call it reads the tail of the transcript and takes the last turn's context (input + cache read + cache write). Once per session it says at 150k: finish soon and hand off with a next-session chip, and at 250k: finish now. Sub-agents are metered on their own transcript. Its markers live in the OS temp folder. |
Lane agents (.claude/agents/, ADR-031) | explorer (small model, read-only, at most 30 lines), queue-runner (middle model, art and audio batches as one script) and reviewer (strongest model, read-only, appends a scorecard to logs/reviews.jsonl). Rule 7 by default. |
Both hooks fail open: bad input or an internal error lets the call through. They are tested in
packages/game-data/tests/guard-hooks.test.mjs. They live in the hooks block of zoen/.claude/settings.json and apply
to sessions opened in zoen/. Claude Code loads them when a session starts. Its settings watcher usually applies an
edit to a running session as well, so start a new session to be sure. To turn them off, delete that hooks block
(or the file). To turn off a lane agent, delete its file.
Tasks finished before this change keep their old, overlapping per-task numbers. The Ledger total is unaffected.
◆4. Faster game development
Accepted on 2026-10-06. In roadmap.json: A is the needs and earlyStart fields, B is M0-09, C is M0-11, D is M2-11
and F is M0-10.
| # | Change | Why it is faster |
|---|---|---|
| A | Start the half that does not need Unity. Rust install and the first crate (M0-03), the server cast-phase state machine with ADR-027 cancels (M2-02), the zone loop and bots (M1-03, M1-04), the mannequin in Blender (M2-09), and the audio and asset validators | Only M0-04, M0-05, M0-07, M1-01, M1-02, M1-05 and the client side of M2 wait for the owner's Unity install. Rust 1.99 and the .NET SDK 10 were installed on 2026-10-06; Unity is not installed yet. |
| B | Engine-free gameplay core. Input arbitration, chain logic, cancel rules and prediction live in plain C# assemblies with no engine references, tested with dotnet test | C1–C18 logic tests then run in seconds without starting the editor, and their logs are small. Needs the free .NET SDK. |
| C | One simulation instead of two (ADR-029). Build zoen-sim (Rust) for the client as a native plugin instead of re-implementing a C# subset | Every movement, cast-phase and cooldown rule is written and tested once. Risk: Unity's Web build links plugins built with the same Emscripten version it ships (3.1.38 since 2023.2; check the pinned editor). One spike task decides; ADR-024 stays the fallback. |
| D | A skill factory after the first polished skill. The 144 skills use 39 shape types that fall into about 8 families (melee arc 25, self area 26, self buff or summon 23, travel 19, targeted 19, projectile 17, ground area 9, line 6). Build each family once as a data-driven runtime, plus one gate runner that iterates every skill and one command that produces the evidence | The roadmap budgets 5–7 hand-built tasks per skill. With a factory a new skill is a data file, shared building blocks and a review. |
| E | Scenes and prefabs from code. The Combat Lab and test scenes are built by an editor script from data | Agents edit C# and JSON, never scene files, and scenes are reproducible. |
| F | A Unity runner that summarises. One wrapper starts batchmode, saves the full log and prints the result, the counts and the first five errors | Editor logs are long; unfiltered, they would be the largest thing in every client task's context. Build it with M0-04. |
| G | Lanes in separate sessions and worktrees. Server, client and content sessions each keep a small context; the integrator merges | Already described in How AI builds Zoen; it needed a clean git state first: the working tree was committed on 2026-10-06 (branch sync/working-tree-20261006). |
| H | Stage the polish. Only the 44 core details of the catalog gate the first skill | The owner judges real feel early, and later details are not built before the systems they sit on are stable. |
◆5. What an agent can and cannot check
The development guide asks agents to watch each animation at full speed, in slow scrub and in loops. An agent does not perceive motion: it reads still frames and numbers. Since 2026-10-06 (ADR-029) the agent's self-review is three things, and watching motion is the owner's part.
- Metrics over every frame: contact frame, planted-foot slide, ground penetration, hand and weapon drift, joint limits, and joint speed spikes (a pop is a spike).
- One contact sheet: the eight views at the four key poses in a single image, plus one strobe strip of the swing.
- Flagged frames only: crops of the frames where a metric failed or came close.
The owner then reviews the clip. This keeps the guide's intent, costs far fewer tokens and does not claim a check that was not made.
◆6. Status of the plan
- Done on 2026-10-06: decisions recorded (ADR-028 to ADR-030), token rules in
CLAUDE.md, Rust and the .NET SDK installed, the working tree committed,CONTINUE_HERE.mdreduced to status and next actions. - Next, one fresh session each: the server lane (M0-03, then M2-02 and M1-03), the content lane (M2-09 mannequin) and the core scaffold (M0-09).
- Owner: install Unity 6.3 LTS or let an agent run the Hub command line, and sign in. Then the client lane starts: M0-04, the runner (M0-10), the time and weather controller (M0-12) and the one-simulation spike (M0-11).
- After the first polished skill: the skill factory (M2-11) before the 18 skills of M4.
- Round 2 (2026-10-07, ADR-031): the rules above were not enforced, and the session that wrote them ran at a 425k median context. The guards (read guard, context meter, lane agents, next-session chips) were built on 2026-10-07 (ENFORCE-20261007-01, section 3); findings are in Development review, round 2.
Source: zoen/docs/process/EFFICIENCY.md · 2,292 words · edit the Markdown, not this page.
