Game guide · source of truth
Tech

1,000 players per map

How one Rust zone process serves 1,000 players and ~4,000 monsters at 20 Hz — spatial hashing, interest management, bandwidth math, prediction, channels, crowd rendering and the load-test plan.

Two separate problems hide behind the 1,000 number. They have separate budgets and separate tests:

  1. Server + network: simulate 1,000 players + ~4,000 monsters, and send each client only what matters.
  2. Client rendering: a phone can't draw 1,000 armoured, animated characters, so it draws the right ones and fakes the rest.
Bandwidth & tick calculator (downstream per client)
12.3 KB/sper client (cap 32)
12.0 MB/sserver egress at 1000 players
14.6 msest. tick (budget p99 15 ms of 50)
Model: 6 B per entity update (bit-packed delta) · interest top-150 by priority · tick model calibrated in M1 bench (placeholder until measured).

◆Server tick (20 Hz = 50 ms budget, target p99 ≤ 15 ms)

System (per tick)WorkBudget (8 vCPU)
Ingest inputs≤ 1,000 command packets, validate, enqueue0.5 ms
Movementintegrate ~5,000 movers vs 1 m walkability grid0.6 ms
Spatial hash rebuild16 m cells (256 × 256 over 4,096 m), entity → cell0.4 ms
Monster AIonly monsters within 60 m of a player are awake (~1,500 peak)1.5 ms
Combat~15 casts/tick, shape sweeps against cells, per-target hit ids0.8 ms
Statuses, cooldowns, regenstruct-of-arrays loops0.4 ms
Interest + priorityper client: gather candidates in 90 m, score, pick top N (parallel by client)3.0 ms
Serialize snapshotsbit-pack deltas per client (parallel)2.5 ms
Network sendtokio writes, batched1.0 ms
Total≈ 11 ms p50, ≤ 15 ms p99

Data layout: struct-of-arrays per component (positions, velocities, hp…), generational entity ids, no heap allocation in the tick (pre-sized buffers). rayon parallelises per-client interest work and per-cell AI.

The numbers the zone server reads (tick budget and targets, the spatial-hash grid, interest radii and set limits, channel caps, the scheduler's catch-up cap) live in game data, packages/game-data/data/netcode.json; this page explains them. services/zone (M1-03) runs the tick on fixed absolute deadlines: after a stall it catches up at most 4 ticks back to back and drops older backlog (proposed). Measured on one thread: logs/evidence/M1-03/tick-bench.json.

◆Interest management (what each client receives)

  • Radius: candidates within 90 m (towns 70 m). Beyond that only party members (minimap) and boss/event beacons.
  • Sets (proposed, M1-03): an entity enters a client's interest set inside the radius and leaves only beyond radius + 5 m (hysteresis, so walking along the edge never flickers). A set holds at most 512 entities (the 9-bit AOI table index); when more qualify, the nearest are kept. Each tick reports the enters and leaves per client.
  • Priority score per entity per client, accumulated every tick: priority += base(type) × falloff(distance) × relevance. Relevance is ×8 for my target, ×6 for anyone attacking me, ×4 for party members, ×10 for bosses and ×3 for monsters casting a telegraph. The top entities that fit the packet budget are sent and their accumulators reset.
  • Rates: near (≤ 30 m) effectively 20 Hz, mid (30–60 m) ~10 Hz, far (60–90 m) ~4 Hz.
  • Events (cast start, hit, death, telegraph) are reliable and sent to everyone in range, independent of the movement priority.

◆Bandwidth math (downstream per client)

ItemBytes
Position (18-bit x/y absolute or 10-bit delta, height from terrain + 8-bit offset), yaw 8 bit, anim state 6 bit4–7
HP/MP change (when changed)2–3
Entity header (index into client's AOI table, 9–10 bit)1–2
Combat event (cast start / hit / death)8–20
SceneVisible entitiesMovementEventsTotal
Quiet field2525 × 6 B × 12 Hz avg ≈ 1.8 KB/s~1 KB/s≈ 3 KB/s
Busy hunting area70≈ 6 KB/s~3 KB/s≈ 9 KB/s
Crowded town (300 nearby)150 sent (cap)≈ 14 KB/s~4 KB/s≈ 18 KB/s
World boss (120 players)150 + boss≈ 14 KB/s~10 KB/s≈ 24 KB/s

Hard cap per client: 32 KB/s. Upstream: 20 input packets/s × ~14 B ≈ 0.3 KB/s.

Built (M1-03B): the wire format is packages/protocol/zoen.schema.json (codec crates/zoen-protocol); bit widths, rates and the proposed values (16 mm position quantum, 10-bit HP permille, 32-frame baseline history) are in netcode.json → snapshot. Each client gets the nearest 150 of its interest set, delta-encoded against the newest frame it acked (a full snapshot when that is older than the history); leaves always go, enters and due changes share the 1,600 B packet. Measured: logs/evidence/M1-03B/snapshot-bench.json (pnpm zone:snapbench). At 1,000 clients the server sends about 10–20 MB/s, which fits on one 1 Gbit NIC with plenty of headroom.

◆Transport and prediction

  • WebSocket (wss), binary frames, with Nagle disabled. WebTransport datagrams are a later experiment, kept only if they measure better on Safari/Chrome.
  • Built (M1-03C): pnpm zone:run -- [--port P] [--bind ADDR] [--seconds S] [--quiet] starts the zone binary (services/zone, src/net/): plain ws:// on loopback port 7310 by default (netcode.json → server, proposed; TLS comes from the gateway later), RFC 6455 written on std (no crates), session and close codes in packages/protocol/zoen.schema.json. Per connection a reader thread (handshake, frame decode, idle timeout) and a writer thread; one bounded queue into the tick and a per-client pool of frame buffers out of it, so the tick never waits on a socket. Clients spawn around the start waypoint's zone centre. Rust test client: zoen_zone::net::client::Client; end-to-end tests in services/zone/tests/server.rs.
  • Load check (M1-03D): pnpm zone:loadcheck -- [--clients N] [--seconds S] [--out PATH] [--quiet] joins 200 headless clients to an in-process server for 10 s (random walk, one input per tick, an ack per snapshot, a ping per second) and fails on a loop p99 over the target, downstream over the 32 KB/s hard cap, stale or dropped inputs, more than 1% skipped frames or dropped ticks, or an unclean close. It also reads the tick thread's CPU time next to its wall time, so waiting for a core on a busy machine shows apart from slow code. Measured: logs/evidence/M1-03D/loadcheck.json.
  • Snapshot pool (M1-04B): per-client snapshots are built and handed off by server.snapshotWorkers threads (4, proposed; the tick thread runs shard 0). Clients sit in stable shards (slot mod workers); the workers read the tick's state, the tick waits for all of them, and the bytes equal the single-threaded path (services/zone/tests/snapshot_pool.rs). Std threads, no rayon: its scoped jobs allocate per tick. 1,000 bots, 80 s: loop p99 12.1 ms (logs/evidence/M1-04B/).
  • Bots (M1-04): pnpm zone:bots -- [--bots N] [--ramp-s R | --steps 100,250,500,1000 [--step-s S]] [--hold-s H] [--workers W] [--snapshot-workers S] [--addr HOST:PORT] [--out PATH] runs the load-test plan below against an in-process server (or a running zone with --addr): bots on poll(2) workers, behaviour mix and pass thresholds from netcode.json → loadTest, loop phase profile in the metrics JSON. --steps is the stepped ramp (each step's new bots connect over the first half of its S s window, the second half holds; default S = loadTest.rampMinutes / steps). The metrics carry per-second series (connected bots, bytes per bot down/up, RTT p50/p99, process RSS; server loop, CPU and phases per 1 s window), the load average each minute and the memory drift over the hold (least-squares RSS slope × hold / mean RSS, limit loadTest.memoryDriftPercent). python3 tools/zone-soak-report.py <metrics.json> --png <chart.png> --summary <summary.txt> draws the dashboard PNG and the pass/fail summary.
  • Own movement is predicted with the client's C# prediction subset, which must match zoen-sim on the shared golden vectors (generated from the Rust crate). Inputs carry sequence numbers; on each snapshot the client rewinds and replays unacknowledged inputs. Errors under 0.05 m are ignored, under 2 m are blended over 100 ms, and larger errors snap.
  • Others are interpolated 100 ms behind (2 snapshots) and extrapolate up to 150 ms on loss.
  • Skills: windup presentation is predicted; damage never is. The attacker's contact sparks may play predicted, but numbers wait for the server.
  • Lag compensation (PvE): for player attacks the server checks monster positions rewound by RTT/2, capped at 150 ms, which favours the attacker. Monster telegraphs lock at cast start so dodges are fair at up to 200 ms RTT.

◆Channels

  • Soft cap 900 (new players are routed elsewhere), hard cap 1,000. When every channel passes 85%, admission opens a new one.
  • Parties join the leader's channel when space allows. Players switch channel at a waypoint, out of combat, with a 10 s cooldown.
  • Bosses and events are per channel. Kharuun spawns in every channel at the scheduled time.

◆Crowd rendering on the client

TierWhoRender pathAnim rate
Ame, party, target, ≤ 20 nearestfull skinned mesh + modular armour, full VFX60 Hz
Bnext 40merged mesh (armour baked to atlas), GPU skinning30 Hz
Cnext 90VAT instanced mesh (one draw per outfit set)15 Hz
Dresthidden (dot on minimap)—

max_visible_players (25/40/60/100/150) caps A+B+C. Nameplates are capped at 40, nearest first. Particle budgets split own 50% / party 20% / others 15% / monsters + telegraphs 15% (telegraphs are never culled).

◆Load-test plan (M1 and M6 gates)

  1. Bots (services/bots, Rust): 1,000 WebSocket clients with scripted behaviours (wander, hunt, crowd the town, cast every 2–6 s, chat).
  2. Ramps: 100 → 250 → 500 → 1,000 over 10 minutes, then hold 10 min (M1) / 60 min (M6).
  3. Pass: tick p99 ≤ 15 ms, no disconnects, avg downstream ≤ 12 KB/s/client, crowded-town clients ≤ 32 KB/s, memory flat (± 5%).
  4. Mixed: 1,000 bots + real browser clients on desktop + a phone. Record the client perf and a 10 s clip in town.
  5. Evidence goes to logs/evidence/M1-04/metrics.json and the histograms as PNG (the M1 soak: logs/evidence/M1-04C/).

Source: zoen/docs/tech/NETCODE_1000.md · 1,705 words · edit the Markdown, not this page.