Game guide · source of truth
AI process

How AI builds Zoen

The exact loop every agent follows — small tasks, tests, logs, token accounting, evidence, reviews — and how the owner stays in control.

The game is built by AI agents (Claude Code as integrator; Codex for images) in small, verified steps. Nothing is "done" because an agent says so. It is done when the task log shows a passing test and the evidence exists. Every step's token cost is recorded, so the owner always knows the running total.

The task loop
start
opens the token window
➜
read
only the docs + data the task needs
➜
implement
one behaviour, ≤ 60 min
➜
test
run by the logger, output saved
➜
evidence
screenshot / 10 s clip / metrics
➜
finish
tokens + tests to the Ledger
300tasks logged (this wiki)
1112tests passed
233tests failed (kept in history)
2460.3Mtokens so far

Live numbers come from logs/tasks.json at build time. See the Ledger.

◆Before game development

Read the game development guide and development-standards.json before character, controls, skill or world work. Record the task's rig/skill/area contract, proposed assumptions and required evidence first. Movement-first is owner direction (ADR-027); numeric tuning is proposed. The mannequin, C1–C18 controls, animation multi-angle self-review and skill/environment gates must pass before expanding content. Wiki previews do not pass engine acceptance tests.

◆1. The task loop

node tools/task-log.mjs start  M2-03 "Mannequin + Horizon Cleave greybox animation" --phase M2
#  … read only the docs/data this task needs, implement the smallest change …
node tools/task-log.mjs test   M2-03 "event times within ±1 frame" -- pnpm --filter @zoen/client test anim-contract
node tools/task-log.mjs note   M2-03 "clip: logs/evidence/M2-03/horizon_cleave.webm"
node tools/task-log.mjs finish M2-03

What each step does:

StepWritesWhy
startlogs/tasks.jsonl start event with UTC timeOpens the token window for this task
testruns the command, saves full output to logs/test-output/<ID>-NN.log, logs pass/fail, exit code, durationTests are run by the logger, so results can't be faked
notefree-text evidence or decisionLeaves a trail for the next session
finishcloses the window and rebuilds logs/tasks.json + logs/token-ledger.jsonTokens per task show up in the wiki Ledger

Rules

  • One behaviour per task. If a task grows past ~60 minutes, finish what passes and open a new task.
  • One task per fresh session (ADR-029): finish, update CONTINUE_HERE.md, start the next task in a new session.
  • Every finish opens the next session (ADR-031). The finishing agent ends by opening the next smallest action as a new desktop-app session (a task chip) with a self-contained prompt: the task id, the files to load and the test that proves it. The owner starts it with one click.
  • Hooks enforce the token rules (tools). The read guard refuses whole reads of the large data files and names the slice to use. The context meter warns once at 150k (finish soon and hand off) and at 250k (finish now).
  • A task needs ≥ 1 passing test before finish --status done. If it can't pass, finish with --status blocked and a note.
  • Failing tests stay in the log. Fix → re-run → the log shows both (red → green), which is the honest history.
  • Visual/feel tasks also need evidence: a screenshot or a 5–10 s clip in logs/evidence/<ID>/.

◆2. How tokens are counted

tools/token-ledger.mjs reads Claude Code's own transcripts (~/.claude/projects/<workspace>/*.jsonl) and counts each assistant message once (message.id), taking the largest value seen for each usage field. It reports:

  • input_tokens + cache_creation_input_tokens + output_tokens: the new work.
  • cache_read_input_tokens: context re-read from cache. It is cheap per token but dominates totals in long sessions.
  • A task's tokens are the messages whose timestamps fall inside its start → finish window, from the sessions that logged its events. Claude Code sessions are attributed automatically (CLAUDE_CODE_SESSION_ID); other agents pass --session. Without a session, concurrent sessions are counted twice (true for most tasks before 2026-10-06). Known offset: the message that issues start can land in the previous window. Start a task in its own short command.
  • Sub-agents. Their transcripts (<session>/subagents/agent-*.jsonl) carry the parent's session id and count in that session. A task that a sub-agent logged (task-log.mjs or verify.mjs calls in its transcript) counts that sub-agent and the helpers it started, not the main session or the other workers, so parallel workers are not double-counted. A task nobody logged from a transcript on disk keeps the session-wide window.
  • finish prints the task's turns and average context. Over 150k tokens per turn it warns: cache re-reads were 69% of all cost in the first three days. Rules and measurements: Speed and token efficiency.

The ledger keeps its history. Claude Code deletes old transcripts, so every rebuild merges the transcripts on disk with the previous logs/token-ledger.json and logs/tasks.json instead of replacing them. A session whose transcript is gone stays and is marked archived; a session in both keeps the larger value; a task's tokens keep the larger value too (tokensArchived, only for the same or a wider window, so closing a task does not keep its open-window figure); totals, by-day and by-model tables, cost and hours never shrink and always add up. Sessions kept from before per-session splits existed sit in the legacy block. Two sessions that hold the same conversation (a resumed copy) are counted once with --counted-in <live>=<archived>; the tools warn when a session is new to the ledger but ended before its last rebuild. Other commands:

  • node tools/token-ledger.mjs --write --fresh (or task-log.mjs report --fresh) drops the history on purpose.
  • node tools/task-log.mjs restore <git-ref> recovers a ledger that an older tool shrank, from its committed copy (one time).
  • Commit logs/token-ledger.json and logs/tasks.json as usual: the history lives in them.

Anything not recorded is reported as not recorded, never as zero. Image generation (Codex) and any paid 3D tools are tracked separately in logs/costs.csv, and only with owner approval.

◆3. Choosing the next task

  1. Read CONTINUE_HERE.md (the previous agent wrote the next smallest action there).
  2. Open the roadmap (node tools/ctx.mjs task M0) → the first task that is not done and whose needs are installed. A task with earlyStart may begin before its milestone's dependency closes.
  3. Too big? Split it with the task template. Good sizes: "add knockback reaction to insectoid set", "pool impact emitters".
  4. Load only what the task needs: the relevant doc section plus the JSON it touches. docs/INDEX.md lists everything. For one skill, monster, zone, roadmap task or catalog area use node tools/ctx.mjs instead of opening the data file.

◆4. Lanes (parallel agents, one owner per resource)

LaneOwnsWrites to
Integrator (main session)repo structure, client app shell, merges, CONTINUE_HEREapps/client-unity, CONTINUE_HERE.md
Data & designpackages/game-data, generators, data testspackages/game-data/**
Serverservices/*, crates/zoen-sim, protocolRust workspace
Characters & animationBlender scripts, rigs, clipsart/source/3d/**, tools/blender/**
Worldterrain generator, kits, streamingapps/client-unity/Assets/Zoen/World/**, tools/world/**
VFX & audioeffect graphs, flipbooks, audio banksapps/client-unity/Assets/Zoen/Fx/**, art/source/vfx/**
UIHUD, windows, wikiapps/client-unity/Assets/Zoen/UI/**, apps/wiki/**
Reviewer (fresh context)reads evidence, never edits codelogs/reviews.jsonl

Only the integrator merges across lanes. Parallel agents work on separate git branches or worktrees.

◆5. Definition of done (per task type)

Task typeMust have
Data changedata tests green; generator re-run if generated; docs updated where numbers are quoted
Server systemunit tests + a deterministic replay test; no unwrap() on server paths; tick budget benchmark when hot
Client featurePlaywright test (renders, no console errors) + screenshot; perf sample if it runs per frame
Skillthe full skill gate
Art assetbudget check (tris/texture size), naming, provenance, in-engine screenshot
UIdesktop 1440×900 + mobile 390×844 screenshots; keyboard + touch paths tested

◆6. Evidence and reviews

  • Screenshots/clips: Playwright (page.screenshot, recordVideo) or the in-game capture hotkey; clips are 5–10 s.
  • On the wiki: everything in logs/evidence/<ID>/ shows on the Ledger (thumbnails, clips, log excerpts; full gallery and notes at /ledger/<ID>/, indexed by tools/evidence-index.mjs on finish). Visual tasks must include a screenshot or clip.
  • Metrics: perf.json (frame times p50/p95/p99, draw calls, triangles, particles), net.json (bytes/s, tick ms).
  • Reviewer agent: after each milestone, a fresh-context agent compares evidence with the docs and art references. It writes a scorecard to logs/reviews.jsonl and lists the three worst defects. Fixes come before new features. The reviewer agent in .claude/agents/ does this on the strongest model and defines the scorecard line. The explorer (searches) and queue-runner (art and audio batches) agents cover the other lanes of ADR-029 rule 7.
  • Owner review: anything visual or "feel" (signature skills, bosses, HUD) needs the owner's yes/no on the clip.

◆7. When things go wrong

  • A tool fails → retry once, then switch to the documented fallback and log a note.
  • Stuck for more than 45 minutes → write what was tried, mark the task blocked, and move to the next task.
  • A spec conflict → don't pick silently. Add an OPEN-* row in Decisions and keep the feature disabled.
  • Never "fix" a failing test by weakening it without a note explaining why the old expectation was wrong.

◆8. What the owner does

  • Approves gates, reviews clips and answers OPEN-* decisions.
  • Generates art from art requests (or asks Codex to).
  • Watches the Ledger page for tokens and test history.
  • Approves any spend. The PHP 5,000 cap covers all new expenses.

Source: zoen/docs/process/AI_WORKFLOW.md · 1,570 words · edit the Markdown, not this page.