How AI builds Zoen
The exact loop every agent follows — small tasks, tests, logs, token accounting, evidence, reviews — and how the owner stays in control.
The game is built by AI agents (Claude Code as integrator; Codex for images) in small, verified steps. Nothing is "done" because an agent says so. It is done when the task log shows a passing test and the evidence exists. Every step's token cost is recorded, so the owner always knows the running total.
◆Before game development
Read the game development guide and development-standards.json before character, controls,
skill or world work. Record the task's rig/skill/area contract, proposed assumptions and required evidence first.
Movement-first is owner direction (ADR-027); numeric tuning is proposed. The mannequin, C1–C18 controls, animation
multi-angle self-review and skill/environment gates must pass before expanding content. Wiki previews do not pass
engine acceptance tests.
◆1. The task loop
node tools/task-log.mjs start M2-03 "Mannequin + Horizon Cleave greybox animation" --phase M2
# … read only the docs/data this task needs, implement the smallest change …
node tools/task-log.mjs test M2-03 "event times within ±1 frame" -- pnpm --filter @zoen/client test anim-contract
node tools/task-log.mjs note M2-03 "clip: logs/evidence/M2-03/horizon_cleave.webm"
node tools/task-log.mjs finish M2-03
What each step does:
| Step | Writes | Why |
|---|---|---|
start | logs/tasks.jsonl start event with UTC time | Opens the token window for this task |
test | runs the command, saves full output to logs/test-output/<ID>-NN.log, logs pass/fail, exit code, duration | Tests are run by the logger, so results can't be faked |
note | free-text evidence or decision | Leaves a trail for the next session |
finish | closes the window and rebuilds logs/tasks.json + logs/token-ledger.json | Tokens per task show up in the wiki Ledger |
Rules
- One behaviour per task. If a task grows past ~60 minutes, finish what passes and open a new task.
- One task per fresh session (ADR-029): finish, update
CONTINUE_HERE.md, start the next task in a new session. - Every finish opens the next session (ADR-031). The finishing agent ends by opening the next smallest action as a new desktop-app session (a task chip) with a self-contained prompt: the task id, the files to load and the test that proves it. The owner starts it with one click.
- Hooks enforce the token rules (tools). The read guard refuses whole reads of the large data files and names the slice to use. The context meter warns once at 150k (finish soon and hand off) and at 250k (finish now).
- A task needs ≥ 1 passing test before
finish --status done. If it can't pass, finish with--status blockedand a note. - Failing tests stay in the log. Fix → re-run → the log shows both (red → green), which is the honest history.
- Visual/feel tasks also need evidence: a screenshot or a 5–10 s clip in
logs/evidence/<ID>/.
◆2. How tokens are counted
tools/token-ledger.mjs reads Claude Code's own transcripts (~/.claude/projects/<workspace>/*.jsonl) and
counts each assistant message once (message.id), taking the largest value seen for each usage field. It reports:
input_tokens+cache_creation_input_tokens+output_tokens: the new work.cache_read_input_tokens: context re-read from cache. It is cheap per token but dominates totals in long sessions.- A task's tokens are the messages whose timestamps fall inside its start → finish window, from the sessions that
logged its events. Claude Code sessions are attributed automatically (
CLAUDE_CODE_SESSION_ID); other agents pass--session. Without a session, concurrent sessions are counted twice (true for most tasks before 2026-10-06). Known offset: the message that issuesstartcan land in the previous window. Start a task in its own short command. - Sub-agents. Their transcripts (
<session>/subagents/agent-*.jsonl) carry the parent's session id and count in that session. A task that a sub-agent logged (task-log.mjsorverify.mjscalls in its transcript) counts that sub-agent and the helpers it started, not the main session or the other workers, so parallel workers are not double-counted. A task nobody logged from a transcript on disk keeps the session-wide window. finishprints the task's turns and average context. Over 150k tokens per turn it warns: cache re-reads were 69% of all cost in the first three days. Rules and measurements: Speed and token efficiency.
The ledger keeps its history. Claude Code deletes old transcripts, so every rebuild merges the transcripts on disk
with the previous logs/token-ledger.json and logs/tasks.json instead of replacing them. A session whose transcript is
gone stays and is marked archived; a session in both keeps the larger value; a task's tokens keep the larger value too
(tokensArchived, only for the same or a wider window, so closing a task does not keep its open-window figure); totals,
by-day and by-model tables, cost and hours never shrink and always add up. Sessions kept from before per-session splits
existed sit in the legacy block. Two sessions that hold the same conversation (a resumed copy) are counted once with
--counted-in <live>=<archived>; the tools warn when a session is new to the ledger but ended before its last rebuild.
Other commands:
node tools/token-ledger.mjs --write --fresh(ortask-log.mjs report --fresh) drops the history on purpose.node tools/task-log.mjs restore <git-ref>recovers a ledger that an older tool shrank, from its committed copy (one time).- Commit
logs/token-ledger.jsonandlogs/tasks.jsonas usual: the history lives in them.
Anything not recorded is reported as not recorded, never as zero. Image generation (Codex) and any paid
3D tools are tracked separately in logs/costs.csv, and only with owner approval.
◆3. Choosing the next task
- Read
CONTINUE_HERE.md(the previous agent wrote the next smallest action there). - Open the roadmap (
node tools/ctx.mjs task M0) → the first task that is not done and whoseneedsare installed. A task withearlyStartmay begin before its milestone's dependency closes. - Too big? Split it with the task template. Good sizes: "add knockback reaction to insectoid set", "pool impact emitters".
- Load only what the task needs: the relevant doc section plus the JSON it touches.
docs/INDEX.mdlists everything. For one skill, monster, zone, roadmap task or catalog area usenode tools/ctx.mjsinstead of opening the data file.
◆4. Lanes (parallel agents, one owner per resource)
| Lane | Owns | Writes to |
|---|---|---|
| Integrator (main session) | repo structure, client app shell, merges, CONTINUE_HERE | apps/client-unity, CONTINUE_HERE.md |
| Data & design | packages/game-data, generators, data tests | packages/game-data/** |
| Server | services/*, crates/zoen-sim, protocol | Rust workspace |
| Characters & animation | Blender scripts, rigs, clips | art/source/3d/**, tools/blender/** |
| World | terrain generator, kits, streaming | apps/client-unity/Assets/Zoen/World/**, tools/world/** |
| VFX & audio | effect graphs, flipbooks, audio banks | apps/client-unity/Assets/Zoen/Fx/**, art/source/vfx/** |
| UI | HUD, windows, wiki | apps/client-unity/Assets/Zoen/UI/**, apps/wiki/** |
| Reviewer (fresh context) | reads evidence, never edits code | logs/reviews.jsonl |
Only the integrator merges across lanes. Parallel agents work on separate git branches or worktrees.
◆5. Definition of done (per task type)
| Task type | Must have |
|---|---|
| Data change | data tests green; generator re-run if generated; docs updated where numbers are quoted |
| Server system | unit tests + a deterministic replay test; no unwrap() on server paths; tick budget benchmark when hot |
| Client feature | Playwright test (renders, no console errors) + screenshot; perf sample if it runs per frame |
| Skill | the full skill gate |
| Art asset | budget check (tris/texture size), naming, provenance, in-engine screenshot |
| UI | desktop 1440×900 + mobile 390×844 screenshots; keyboard + touch paths tested |
◆6. Evidence and reviews
- Screenshots/clips: Playwright (
page.screenshot,recordVideo) or the in-game capture hotkey; clips are 5–10 s. - On the wiki: everything in
logs/evidence/<ID>/shows on the Ledger (thumbnails, clips, log excerpts; full gallery and notes at/ledger/<ID>/, indexed bytools/evidence-index.mjsonfinish). Visual tasks must include a screenshot or clip. - Metrics:
perf.json(frame times p50/p95/p99, draw calls, triangles, particles),net.json(bytes/s, tick ms). - Reviewer agent: after each milestone, a fresh-context agent compares evidence with the docs and art references.
It writes a scorecard to
logs/reviews.jsonland lists the three worst defects. Fixes come before new features. Therevieweragent in.claude/agents/does this on the strongest model and defines the scorecard line. Theexplorer(searches) andqueue-runner(art and audio batches) agents cover the other lanes of ADR-029 rule 7. - Owner review: anything visual or "feel" (signature skills, bosses, HUD) needs the owner's yes/no on the clip.
◆7. When things go wrong
- A tool fails → retry once, then switch to the documented fallback and log a note.
- Stuck for more than 45 minutes → write what was tried, mark the task
blocked, and move to the next task. - A spec conflict → don't pick silently. Add an
OPEN-*row in Decisions and keep the feature disabled. - Never "fix" a failing test by weakening it without a note explaining why the old expectation was wrong.
◆8. What the owner does
- Approves gates, reviews clips and answers
OPEN-*decisions. - Generates art from art requests (or asks Codex to).
- Watches the Ledger page for tokens and test history.
- Approves any spend. The PHP 5,000 cap covers all new expenses.
Source: zoen/docs/process/AI_WORKFLOW.md · 1,570 words · edit the Markdown, not this page.
