Technical architecture

How Halvesper runs: one world on desktop, web and mobile, 1–5k CCU at launch, built on what the repo already has. Canon (00-canon.md) is binding; docs/platform.md and docs/plan.md are the earlier decisions this document extends. Lore, systems and art are other documents; this one covers only what they run on.

Every recommendation is labeled: - Decided: already in the code or in canon/platform, or a choice with no real alternative. Changing it is a deliberate canon change. - Proposal: the recommended path. Foster can overrule it without a spike. - Needs a spike: the answer depends on a measurement we don’t have yet. The spike and the result that decides it are in §13.

1. What exists today (2026-10-03)

Piece Where State
Pure deterministic sim, 60 Hz, tick(frames) keyed by player id game/sim/sim.gd, zone.gd Works. RefCounted only, no scene tree; lint enforces it.
Zones from text maps, 8-unit tiles (32 reference px), one-way platforms, ropes, ladders, portals game/content/maps/*.map.json, footholds.gd Lantern Town 1920×420 units, Mossy Field 1920×320, Footprint Ponds 2080×380, the Old Tollbridge 1700×360. Footholds since S3. The tile rows in *.art.json still load but are no longer drawn: game/render/ground.gd paints surface strips and ropes from the footholds and climbables, and the map’s props place painted art.
Authoritative server, transport-free game/net/server.gd Per-client input buffer (max 8 ticks), 20 Hz snapshots of the client’s own zone, var_to_bytes dictionaries.
Client prediction and interpolation game/net/client.gd Predicts own movement, replays unacked input on each snapshot; others drawn 6 ticks (100 ms) behind.
Gameplay light game/sim/light.gd, game/render/light_overlay.gd Analytic circles (03 §6.1): every player’s rank-1 Open lantern plus a map’s static lights, spatially hashed (128-unit cells) and rebuilt each tick; ids in snapshots, own lantern predicted. Puddlings, Dimwings and Afternooners are Lampghosts (Shroud/Lit/Seared). The renderer darkens by the map’s ambient tier and lifts the same circles.
Lantern trip loop game/sim/lantern.gd, mote.gd, game/ui/lantern_gauge.gd Rough (03 §7.2): motes released by Lampghosts and bosses and caught; on a map without carry_motes (every field and hold) they bank as caught, and on one with it fuel burns by tier and load and motes are carried, leaked, spilled and banked at a hearth or on arrival on a non-carrying map; motes are entities in snapshots, fuel and motes are ME fields, the Sunline is a snapshot field kept in server memory.
Headless WebSocket server server/server.gd One process holds every zone. Prints SERVER_METRICS every 10 s.
Monster look game/render/enemy_view.gd, lampghost.gdshader, tools/art/monsters.py Sprites generated locally with mflux (prompts and seeds in tools/art/monsters.json), 4 texels per unit. Lampghosts are painted pale (red is lightness, blue the stolen ember’s field) and drawn above the light overlay in their kind’s coldlight tint, a little see-through with drifting specks, a pale rim and halo, and an ember that widens and brightens with the fire they hold (their health); dimmer when Shrouded. The Cinder Eye is painted. Single sprites with procedural squash, not cutout rigs yet.
Monster AI as data game/sim/creatures.gd (ai), behaviours.gd, brains.gd, native/src/monster_ai.cpp Four archetypes (hop, fly, walk, hover-dash) with per-kind numbers, run in one C++ call per zone; the GDScript path is bit-identical (tests/test_behaviours.gd) and runs on the web. A kind may name an event hook, a godot-sandbox v0.60 SafeGDScript run on the server only (4×2^20 instructions and 16 MB per call, answers clamped; the server refuses to start if one fails), §5.3; no built-in kind has one.
CI tools/gate.sh, .github/workflows/gate.yml Gate on macmini4 ci: import, lint, test, scenarios, online (server plus two headless clients over real sockets, each of which must see the other walk), shots. Spikes fail on parity or script failures, never on timing. Nightly seed sweeps and a web boot smoke; the playable web build is the one deploy/deploy.sh serves.
Deploy deploy/ fdhproperty OVH VPS (4 vCPU, 7 GB): nginx terminates wss at /ws, systemd delve-server capped at 1 GB and 1 CPU, Godot editor binary run with --headless. Postgres 17 on localhost and delve-worldd beside it; deploy/host.sh makes databases and keeps secrets in /opt/delve/env (root, 0600).
worldd v1 worldd/ (Go, pgx, stdlib HTTP), game/net/ticket.gd, character.gd, server/worldd.gd Building (Phase 1 item 1): characters per account, join tickets, whole-character saves with item uids and the ledger; the game server checks tickets and loads and saves through it. See the join flow in §3.

Measured so far (these are the only real numbers; everything else below is a target to measure):

Metric Value Source
Server tick, idle ~0.05 ms SERVER_METRICS on fdhproperty
Tick with 2 players walking 0.26 ms PR #13
Bandwidth per player (2 players, few monsters) ~1.6 KB/s down PR #13
Client RAM, desktop 41 MB PR #13
Monster AI, old path ~0.14 ms per monster per tick (15 monsters = 2.2 ms) PR #15 on fdhproperty
Monster AI, batched (S1) 29–37 µs per monster-tick on fdhproperty (sandbox think 4.5–5.5, of which the call is 0.31–0.40; host movement 25–31); 6.95 µs on macmini4 S1, PR #17, 2026-10-03
Body.move (GDScript host physics) ~27 µs per body-tick on fdhproperty S1, PR #17
Monster AI as data (S1 reopened) think 0.85–0.90 µs per monster-tick native (sandbox 5.8–6.9, GDScript data 4.4–4.6), plus native movement 1.5–1.8: ~2.6 µs per monster all in spike_s1, PR #26, 3 runs on fdhproperty
Player cost implied by the 2-player tick ~100 µs per player per tick (includes var_to_bytes snapshots) PR #13
Prediction corrections at 100 ms each way 0 tests/test_net.gd

S1 found sandboxed scripts cheap enough for all monster AI, but once S1b made movement native the sandbox think was the larger cost, so S1 was reopened: monster behaviour is data run by native code, and the sandbox is for creator hooks (§5.3). The real cost was GDScript host physics. Body.move costs ~27 µs per body-tick on the VPS, and players use the same Body.move, so §5.3’s ~3 µs per player cannot hold in GDScript. At ~100 µs per player, a 6 ms tick holds ~60 players per process. Moving body movement and collision into a C++ GDExtension shared by server and client (spike S1b, Phase 0) is the plan; §5.3 has the fallback ladder.

2. Goals, non-goals and targets

Goals

  1. One world, one account, on desktop, web and mobile (decided, canon).
  2. Server-authoritative everything: movement, combat, drops, XP, economy (decided).
  3. Feel first: own-character movement answers in the same frame on every platform, and bots in CI keep it that way (decided, canon pillar 5).
  4. Never lose or duplicate a traded item (decided here; §7).
  5. Cheap to run: bare metal, not managed game hosting (decided in platform.md).
  6. Run by a very small team: few moving parts, each one boring (proposal; this drives most build-vs-adopt calls below).

Non-goals (for launch)

Quantified targets

Target Value Status
CCU, launch design point 5,000 peak, one world Decided (canon 1–5k)
CCU, headroom path 20,000 with linear hardware, no redesign Proposal
Players per map instance 150 cap Decided (canon)
Visible players per client 60 nearest (mobile default 30) Proposal
World boss attackers 60 per boss instance Decided (canon)
Players per channel process 400–600 target; ~60 measured in GDScript today Needs S1b and S2
Reference viewport 1920×1080 Decided (canon)
Sim tick 60 Hz on server and client Decided (plan “Kept”)
Snapshot rate 20 Hz (every 3 ticks) Decided (Protocol.SNAPSHOT_EVERY)
Input send rate 30 Hz (2 frames per packet) with redundancy Proposal (today 60 Hz)
Interpolation delay 100 ms (6 ticks) Decided (INTERPOLATION_TICKS)
Server tick time per channel process mean ≤ 6 ms, p99 ≤ 12 ms of 16.7 Proposal
RTT p75 ≤ 80 ms, p95 ≤ 150 ms in the US from one central DC Proposal
Playable RTT ≤ 250 ms with no gameplay errors, only visible lag Proposal
Downstream per player median ≤ 4 KB/s, p99 ≤ 12 KB/s (crowded town) Proposal
Upstream per player ≤ 2 KB/s Proposal
Map transfer (portal, same channel) ≤ 300 ms server side Proposal
Channel/host transfer ≤ 1.5 s p95 including reconnect Proposal
Server cost ≤ $0.25 per peak CCU per month at 2k, ≤ $0.15 at 5k Proposal (§12)
Persistence Trades, market, mail, crafting, NPC buys: committed before acknowledged (RPO 0). Other progress (XP, common drops, position): RPO ≤ 30 s Proposal
Database disaster RPO ≤ 1 min (WAL archive), RTO ≤ 1 h Proposal
Uptime 99.5% monthly excluding announced maintenance at launch; 99.9% from year two Proposal
Maintenance One weekly window ≤ 60 min until rolling deploys are proven; content-only patches with no downtime Proposal

Latency budget (one input, US player, 80 ms RTT)

Step ms
Own movement shown (predicted locally) 0–17 (next frame)
Input reaches server (RTT/2) + buffer (≤ 1 tick normal) 40 + 17
Server applies, next snapshot (≤ 3 ticks) ≤ 50
Snapshot reaches client 40
Server-confirmed result visible (damage number, drop) ~150 typical, ~250 p95
Other players seen their RTT/2 + 100 ms interpolation + yours/2 ≈ 180 ms behind real time

3. Topology

Diagram

                          players: web (WASM), desktop, iOS, Android
                                          |
            HTTPS (client, packs)         |  WSS (login, social)        WSS (gameplay)
     +--------------------------+         |                                 |
     | Cloudflare CDN + R2      |         v                                 v
     | web client, asset packs, |   +-----------+   +-------------+   +----------------------+
     | patch manifests          |   |  nginx    |-->|  Nakama     |   | game host 1..N       |
     +--------------------------+   | (svc host)|   |  accounts,  |   |  nginx: wss /c/<n>   |
                                    |           |   |  platform   |   |   -> 127.0.0.1:port  |
                                    |           |   |  logins,    |   |  channel procs (Godot|
                                    |           |-->|  friends,   |   |  --headless, 1/core) |
                                    +-----------+   |  groups,    |   |  spawner (Go agent)  |
                                         |          |  chat, IAP  |   +----------+-----------+
                                         v          +------+------+              |
                                    +-----------+          |          control plane: JSON over
                                    |  worldd   |<---------+----------WebSocket on WireGuard
                                    |  (Go)     |                                |
                                    |  directory, tickets, transfer, parties,    |
                                    |  market, mail, escrow, GM API, events      |
                                    +-----+-----+<-------------------------------+
                                          |
                               +----------v----------+       +-----------------------+
                               | PostgreSQL 17       |------>| replica (from beta)   |
                               | nakama db, world db |  WAL  | pgBackRest -> object  |
                               +---------------------+       | storage (PITR)        |
                                                             +-----------------------+
     observability: Vector (logs + SERVER_METRICS) -> Prometheus/Loki -> Grafana, alerts

In three sentences: players log in through Nakama and get a channel list and a signed join ticket from worldd, a small Go service that owns the world directory, transfers and every cross-character economy operation in PostgreSQL. Gameplay runs in channel processes, each a headless Godot server running the existing pure sim with every map of one channel, one process per core on cheap US bare metal, with nginx terminating wss on each host. Portals inside a channel stay in-process (as today); changing channel or host is a ticketed reconnect that saves the character first.

Services: build or adopt

Service Choice Status Why
Gateway / TLS nginx on every host Decided (already deployed) Proven in deploy/nginx-delve.conf; terminates wss, rate-limits, hides ports. No custom gateway process.
Accounts and auth, platform logins Adopt Nakama (Apache 2.0) Decided 2026-10-04 (Foster), building; S6’s platform-login half still open Device, email, Apple, Google, Steam, custom auth and account linking out of the box; IAP receipt validation for Apple and Google. Building these is weeks of security-sensitive work. Built so far: Nakama 3.41.0 from its release binary (checksum pinned in deploy/deploy.sh), own database delve_nakama on the host’s Postgres 17, systemd delve-nakama, nginx /nakama/ to its HTTP API and socket on localhost (console localhost only), secrets in /opt/delve/env. The client signs in with guest (device id) or email and password over plain HTTPRequest against Nakama’s REST API (game/net/account.gd), not the official Godot addon: four endpoints don’t justify an autoload and a vendored SDK, and a transport-free state machine is unit-testable.
Friends, guild membership, social graph Nakama (friends, groups) Proposal Groups map onto guilds (membership, roles); guild game state (bank, level) lives in worldd.
Chat Nakama channels carrying preset-phrase ids Proposal Preset-only at launch means no moderation stack; Nakama gives rooms, groups and presence. Map-local speech bubbles go through the channel process instead (§5).
Party worldd (state) + channel process (shared XP) Building: worldd’s party tables and internal API (worldd/party.go), polled each second by the one game host with its online characters; the channel process shares kills by party on the map (03-systems.md §5.5). TODO: per-host presence and a LISTEN/NOTIFY push once there is more than one host Party members can be on different channels; worldd is the cross-process truth. Nakama’s realtime parties are matchmaking-shaped and don’t own game rules.
World directory, channel selection Build: worldd Building: v1 is characters, saves, item uids, the ledger and join tickets for one channel (worldd/) Small, game-specific: which channel process hosts what, load per channel, join tickets, transfers, titan positions.
Map/channel servers Build: existing Godot server Decided game/net/server.gd + server/server.gd, extended (§5).
Transfers (maps, channels, holds) Channel process + worldd Proposal §4.
Market / auction worldd + Postgres Proposal Must be transactional across characters; §7.
Mail worldd + Postgres Proposal Same escrow path as the market.
Persistence Adopt PostgreSQL 17 Decided Transactions, constraints and row locks are the anti-dupe design. Nakama already needs Postgres (or CockroachDB), so it is one database server, two databases.
Cache / presence across hosts Valkey Rejected for launch worldd keeps presence in memory and rebuilds it from channel re-registration after a restart. Adopt Valkey only when worldd runs more than one replica.
Event bus Postgres LISTEN/NOTIFY, fanned out by worldd Proposal World events (Sunline Strides, Act changes, world boss spawns, Kneelings, market fills) are written as rows in world_events, then NOTIFY carries only the row id. worldd is the one listener and pushes events to channel processes over its existing control links; consumers catch up by id after a reconnect, so nothing is lost when a listener drops. No new service. NOTIFY payloads stay under 8 KB and are not durable, which is why the row is the truth.
Message bus NATS Rejected for launch Godot has no client; worldd plus LISTEN/NOTIFY covers a few hundred events per second. Revisit past ~20k CCU.
Process orchestration Build: spawner (Go agent per host, systemd transient units) Proposal platform.md already calls for “a small spawner”. systemd gives CPU and memory caps and restarts; Kubernetes would be more to operate than everything else combined.
Analytics Postgres partitioned events table → ClickHouse later Proposal Fine to ~50M events/day; move when queries hurt.
Admin / GM tools worldd admin API + small web UI; Nakama console for accounts Proposal Every GM action is audited (§9).
Observability Prometheus, Loki, Grafana, Vector, Alertmanager (self-hosted) Proposal Free, standard, and Vector can turn the existing SERVER_METRICS lines into metrics without touching the server.
CDN and downloads Cloudflare + R2 Proposal No egress fees; the plan already found Cloudflare Pages caps files at 25 MiB, R2 does not.
Hosting OVH US bare metal (Rise line), not Hetzner, for launch Proposal (changes platform.md) See below.

Join flow (as built, worldd v1)

  1. The client logs in to Nakama (device id for a guest, or email and password) and holds a Nakama session JWT and refresh token, stored on the device and refreshed when under 5 minutes are left or when worldd answers 401 (once); the login screen and character select are game/ui/login_screen.gd over game/net/account.gd.
  2. With it as Authorization: Bearer, the client calls worldd (same origin, nginx /worldd/): GET /v1/characters, POST /v1/characters {name, class} (class warden, kindler, ranger or shade; names 3–12 letters after trimming, with at most one space, apostrophe or hyphen between letters, unique ignoring case, 400 when invalid and 409 when taken; 6 per account), and POST /v1/join {character_id}. worldd checks the JWT with Nakama’s session key (HS256), creates the account row on first sight (id = Nakama user id), and answers the join with {character, ticket}.
  3. The ticket is base64url(JSON {account_id, character_id, channel, expires_unix}). base64url(HMAC-SHA256), valid 60 s, signed with a secret shared with the game server. The client sends it in its hello ({t, v, ticket}).
  4. The game server checks it offline (game/net/ticket.gd; once per ticket, and one connection per character per process), then loads the character with GET /v1/internal/characters/{id} (service token) and only then welcomes the client. A character never saved gets the sim’s new kit in town.
  5. It saves the whole character with PUT /v1/internal/characters/{id} on a hearth bank, on a map change, every 60 s and on disconnect, one save in flight per character. Each save carries save_seq + 1 and a reason; worldd rejects one not above the stored save_seq (409). A 409 means another session wrote the character: the server drops its copy unsaved and sends KICK {reason}, and the client goes back to login with that message. On stop (every deploy) systemd’s ExecStop creates the server’s stop file, since Godot does not hand SIGTERM to scripts; the server stops accepting, saves everyone with reason shutdown, waits up to 5 s for worldd and exits (TimeoutStopSec=20).
  6. With no worldd configured (DELVE_WORLDD_URL unset: dev, CI, scenarios, bots) the server takes anonymous hellos as before and nothing persists. The client logs in only when given both Nakama and worldd (--nakama=/--worldd=, or ?nakama=/nakama&worldd=/worldd, which nginx adds to a bare visit) and not --offline.
  7. When the socket closes on a logged-in client it goes back to character select and says why: the server’s KICK reason, a refused ticket (expired, or the character is already playing) when it was never welcomed, or a lost connection. A new join takes a new ticket.

Why worldd is its own Go service and not Nakama runtime code. Nakama’s Go runtime plugins must match Nakama’s exact Go and Nakama versions, which couples the economy to Nakama’s upgrade cycle. A separate ~5–10k-line Go service with pgx, verifying Nakama’s JWT session tokens, is easier to test, deploy and replace. If S6 shows Nakama is a poor fit, only accounts and social need replacing; worldd and the game don’t move.

Why US hosting changes from Hetzner to OVH. platform.md assumed Hetzner bare metal. Two facts changed: Hetzner sells dedicated servers only in Germany and Finland (its US sites are cloud-only) [1], and its 2026 repricing put an AX42 at €97.30/month [2]. US cloud instances there include only 1 TB of traffic [3]. OVH’s US Rise line (Vint Hill VA and Hillsboro OR) has a Ryzen 7 9700X, 8 cores, 64 GB, unmetered 1 Gbps for $77/month [4], and fdhproperty is already an OVH account. A single central US DC (Vint Hill) keeps one world simple; Hillsboro is the failover/West option. Hetzner Germany remains the obvious EU home later.

No private network on Rise-S. The Rise-S plan has no private (vRack) bandwidth, so every link between hosts (worldd control plane, Postgres streaming replication, pgBackRest, metrics) runs over WireGuard on the public NIC. That traffic counts against the 1 Gbps port: budget ~50 Mbps on the DB primary for WAL streaming and backups at launch and alert at 200 Mbps. Postgres listens only on the WireGuard interface. Moving the DB pair to a plan with vRack is a later option, not a launch need.

4. Worlds, channels, maps and transfers

Units (proposal, sized by S2)

Unit Meaning Count at 5k CCU
World The one world. One worldd, one database. 1
Channel A full copy of every public map, run by one process (canon). Players pick one; friends meet by picking the same one. 10–14 at the 400–600 target; set by S1b/S2
Channel process One Godot --headless process hosting all public maps of one channel (what server.gd already is: one Sim with every zone). = channels
Map instance One Zone inside a process. Ticks only while someone is in it (already true: sim.tick skips empty zones and zone.remove resets an emptied map). ~100 maps × channels, mostly idle
Instance process A process hosting private instances: party dungeons, the raid, boss rooms. Pooled; one process holds many instances. 2–4

Why channel-per-process and not map-per-process: in-channel portals stay what they are today (an in-memory move_to), so 90%+ of transfers cost no reconnect and no save; idle maps cost nothing; and a host runs one process per core with no cross-process traffic inside a channel. The fallback, if S2 shows a channel can’t hold ~400 players in the tick budget, is to split a channel by region (one process per region-channel pair) and make inter-region portals ticketed transfers, which the transfer path below supports anyway.

A crowded map (town at peak) is capped at 150. A portal into a full map offers “this map is crowded: switch channel?” (proposal; the systems writer may prefer overflow instances).

Density by CCU. A world that feels empty fails faster than one that lags. worldd opens and merges channels to hold density, not just to absorb load:

CCU Open channels Target per channel Busiest field map Town at peak
≤ 300 1–2 150–300 10–20 60–100
2,000 4–6 350–500 15–30 100–150
5,000 10–14 400–500 15–30 120–150

World bosses run in an instance process, not a channel: up to 60 attackers drawn from every channel, entered through a ticketed transfer from the boss’s field.

World singletons

Some things exist once per world, not once per channel: the titans and their routes, Wend’s Absence, Act state, the Pale Tenant’s resolution, the Sunline. These are world-singleton entities: worldd owns their state in Postgres, each channel process renders a read-only copy, and players’ contributions (motes banked, lights lit, boss damage) are reported up to worldd in aggregates every few seconds. Changes that alter the world (Sunline Strides, Act changes, route parameters) take effect only at a Turning boundary (the weekly reset), announced in advance on the event bus, so every process switches on the same tick. Inside an Act the only live changes are event-scoped (a world boss spawning, a Kneeling).

A one-per-world fight (Wend’s Absence, the Pale Tenant, Morrow’s birth) runs as many 60-attacker instances against one shared HP pool: each instance reports its damage to worldd every 1 s, worldd sums it and broadcasts the pool, and when it empties every instance gets the result on the same tick (03-systems.md §15.2).

Transfer kinds

Kind Path Cost
Portal within a channel Sim.take_portal → move_to, as now. Client gets a snapshot with a new zone and rebuilds the view (client._forget_zone). One tick + map load on the client
Channel change, or a portal into another process (instance, split region) Ticketed handoff (below) ~0.5–1.5 s behind a fade
Login Nakama auth → worldd join ticket → connect ~1–2 s

Ticketed handoff (proposal): 1. Source process freezes the character (no input applied), serializes it and asks worldd to save with its lease epoch (§7.3). 2. worldd commits the save, bumps the epoch to the target process, picks the target (channel, map, entry point) and returns a ticket: {character, target, epoch, expires (+30 s), nonce}, HMAC-signed with a key the channel processes share. 3. Source sends TRANSFER {url, ticket} and drops the character from its sim. 4. Client opens wss://gN.<domain>/c/<k>, sends HELLO {v, ticket}. Target verifies the HMAC offline, asks worldd for the character (one row read plus items), places it. 5. A ticket is single-use (worldd records the nonce); an expired or replayed ticket is refused, and the client falls back to login.

This replaces today’s anonymous HELLO {v} and per-connection throwaway characters.

Titan holds that move

The holds are towns on walking titans (canon). Architecturally a hold is an ordinary set of maps, hosted in every channel like any other; what moves is which region the hold is above. Proposal:

This keeps the “world map never quite the same twice” pillar entirely in data plus a clock. Status: proposal; the systems writer owns the routes.

5. Simulation and netcode evolution

5.1 What stays, changes, is replaced

Area Stays Changes Replaced
Sim core Pure RefCounted sim, tick(frames) keyed by id, preload not class_name, seeds passed in, deterministic Per-zone tick split out so a process can tick zones in any order and measure each —
Collision Axis-separated AABB moves, no Godot physics Grid → footholds (segments, slopes) 8 px tile grid as the collision source
Maps Static data loaded by the sim, diffable, portals by digit/id Authored in the Godot editor with a plugin, exported to data Character-per-tile ROWS scripts (keep as a test format)
Monsters Behaviour as data (archetype plus numbers) run by native sim code, think at 12 Hz, movement integrated natively; sandboxed SafeGDScript only for creator hooks, clamped, server refuses broken scripts Event hooks (spawn, hit, phase, death) in place of a per-tick scripted think (done, §5.3); per-zone monster state held in C++ One Dictionary call per monster per tick; a sandbox call per kind per think
Protocol Server authority, input redundancy, ack-based replay, interpolation 100 ms Binary packing, quantization, deltas vs acked baseline, AOI, staggered snapshots, tickets var_to_bytes dictionaries
Server process server/server.gd loop, SERVER_METRICS Many processes per host, spawner, worldd link, dedicated-server export Single world process with every player

5.2 Hand-built platform maps (needs a spike: S3)

MapleStory’s collision is footholds: line segments, often sloped, that bodies stand on, plus walls, ropes and ladders as vertical segments. Its maps are painted backgrounds and tile props, not a collision grid. The canon (hand-built maps, HD painted art, slopes implied by the style) points the same way. Proposal:

Decides: S3 shows slopes and one-way drop-through with zero prediction corrections at 100 ms RTT and no movement regression in the zone-tour bot. If foothold physics proves painful in the pure sim, the fallback is the current grid at 4 px with slope tiles.

Mover result (PR #28): game/sim/footholds.gd is the collision behind the zone’s seam, with NativeFootholds in C++ matching it bit for bit. Zones still trace their tile ROWS through Footholds.from_grid. The mover is stateless: a grounded body snaps to the nearest foothold within its slope reach, and an airborne one lands on the highest foothold it crosses, rising or falling: a body knocked up a slope faster than it rises lands on the slope instead of tunnelling under it (the grind bot found it on Mossy Field). No body may end a tick under the ground (Footholds.sunk, held by Zone.sunk in the fight scenarios). Prediction and the protocol are unchanged. The spatial hash uses 32-unit columns. Ground is solid, platforms are one-way. Three things change from the grid mover: feet are the box’s bottom centre, there are no ceilings, and there is no step-up (blocked walkers jump). The zone-tour and slime-fight sweeps pass 64 of 64 seeds, and a slope, platform jump and drop-through at 100 ms each way need no corrections (test_net.gd).

Map files (PR #29): maps are game/content/maps/<id>.map.json, read and validated by map_file.gd and registered by id in maps.gd (town 1, field 2, ponds 3, tollbridge 4). The text ROWS were converted once, to the same footholds the tracer gave. The rows survive as <id>.art.json; since the look pass the renderer no longer draws them (they still load, ids append-only, for the scenario that carves the grid) and they are retired as art; maps from the Mossway rebuild on (field, ponds, tollbridge) have none. mapkit (PR #30) round-trips every map. Mossy Field is moss mounds under two tiers of decks, drawn by game/render/ground.gd from its footholds. Still to do: time an editor round trip by hand. Maps may also name an ambient tier, list static lights as [x, y, radius, intensity], a hearth rect [x, y, w, h] where motes bank and lanterns refuel (town only), set carry_motes (true only for the Deep and the strata), and place props, all optional and appended to the key order. A prop is either art ({art, feet, width, flip}: assets/world/<art>.png, bottom centre at feet; render only) or a prop def placed by name ({def, feet, flip}): assets/world/<def>.json beside its art gives the art, its ref-px size and anchor, and its footholds and climbables in the art’s own ref px, and map_file.gd place adds them to the zone’s collision, mirrored when flipped, so a building’s porch, balcony, roof and ladders are where it is painted (props.gd draws the art behind the ground, or over it for a def with front, or over the characters for a def with over, such as pond water they wade in; ground.gd paints only the map’s own footholds, and every climbable). Footholds stay the sim’s only collision (town is Hearth with its hearth, Mossy Field is Dusk with one lamp). The grid mover is gone (PR #32): body.gd is state only, and S1b’s fuzz and kernel timing run on NativeFootholds.

Units: 1 sim unit = 4 ref px at 1920×1080; tiles.gd’s SIZE := 8 stays, so an 8-unit tile is 32 ref px, used only as a map-authoring snap (canon). project.godot is 1920×1080 with linear filtering and mipmaps, and the camera’s zoom 4 shows 480×270 units.

The Deep (08-scale.md §4) is assembled from hand-built pieces by seed: each piece is a small .map.json with typed edge sockets, and a deterministic assembler joins pieces into a map at instance creation. The assembler runs in the sim (server and client get the same seed), and the map lint runs on every piece and on 1,000 seeded assemblies in the Gate.

5.3 Entities per map and tick budget (S1 measured; S1b and S2 decide)

Per map instance (proposal): 150 players, 60 monsters (a busy field 25–40; boss maps fewer, bigger), 300 ground drops, 150 projectiles/effects.

Per channel process at 500 players, 60 Hz, on a Ryzen 9700X core. The “target” column is the budget; the “measured” column is what GDScript does today.

Work Target ms per tick (target) Measured (fdhproperty)
Player movement and actions 500 × ≤ 3 µs 1.5 ~100 µs per player including snapshots; Body.move alone ~27 µs
Monsters ~400 awake × ≤ 5 µs 2.0 ~2.6 µs: native data think 0.85–0.90, native movement 1.5–1.8 (passes)
Light queries ≤ 0.5 µs per body per light test, spatial-hashed 0.5 2.2 µs per query in GDScript with 150 lanterns (128-unit hash), 137 µs to rehash per tick (spike_s1b light_us, macmini4)
Combat, drops, spawns 1.0
Snapshot encode 500/3 clients per tick (staggered) × ~10 µs 1.7 Today var_to_bytes; protocol v2 replaces it
Socket I/O 0.8
Total ~7.5 ms mean, of 16.7 Far over in GDScript

Monsters (S1 reopened 2026-10-03). Monster behaviour is data run by native sim code, as skills are (D7): each kind’s ai names an archetype (hop, fly, walk, hover-dash) and its numbers, and one C++ call per zone thinks every due monster. The GDScript path in behaviours.gd is the reference and matches it bit for bit, so the web build plays offline with real monsters. New archetypes are engine work; new monsters are data. Sandbox scripts are for creator hooks: event handlers (spawn, hit, phase change, death) that adjust a data behaviour, not a per-tick think.

Monster hooks (decided). A kind’s def may name a hook, a SafeGDScript with one entry point, hook(e: Dictionary) -> Dictionary, contract in game/sim/brains.gd:

The per-tick whole-script think is retired from the game; the S1 bench keeps it (tools/bench/s1_brains.gd, tests/test_brains.gd) to measure the contract below. An example hook lives with the tests (tests/hooks/cinder_eye.sgd), not in content.

S1 (run 2026-10-03, PR #17) measured that contract: think_batch(dt, m: Array) -> Array, one call per monster kind, 14 values in and 6 out per monster, as plain Array (in godot-sandbox v0.60, indexing a PackedFloat64Array costs ~65k instructions per access). Monsters think at 12 Hz, staggered by list position; the host integrates movement every tick; the per-call budget is 4 units. At 4.5–5.5 µs per monster-tick it passed S1’s line, but after S1b put movement at ~2 µs it was the largest monster cost, hence the reopening. Data run natively thinks in 0.85–0.90 µs per monster-tick on the VPS (PR #26), about 7× less.

Bodies are the bottleneck. Players and monsters share Body.move, which costs ~27 µs per body-tick in GDScript on the VPS. Fallback ladder, in order, each step tried only if the one before misses the budget: 1. S1b (Phase 0): move Body.move and collision into a C++ GDExtension that is the shared sim core, used by server and client so prediction stays exact. Target ≤ 2 µs per body-tick on the VPS. Golden-vector tests run the GDScript and C++ paths on the same inputs and require identical output (positions are integers in sim units, so bit-exact parity is achievable). The web build needs a custom export template with dlink enabled so it can load the extension; that template is part of S1b. Result (PR #19): 1.7–2.0 µs on the VPS; parity is bit-exact over 2.6M body-ticks with float32 positions, as both paths do the same float operations. The C++ kernel itself is 0.06–0.08 µs, so nearly all of what is left is the GDScript-to-C++ call. The next gain is state held in C++ per zone, not a faster kernel. The grid mover is frozen until S3; the foothold mover replaces it in C++ with the same parity harness. 2. Profile and trim the remaining GDScript (snapshot packing moves to the same extension; protocol v2 removes var_to_bytes). 3. Split a channel by region (one process per region-channel pair), with inter-region portals as ticketed transfers. More processes, same players per core. 4. Drop the sim to 30 Hz with 60 Hz client interpolation. Halves cost; costs feel, so it is the last resort.

Light in the tick. Light radius, mote carry and Lampghost tests are sim queries against the same spatial hash as footholds, inside the budget above. Lampghosts that are hidden by darkness are filtered per client: the server sends a hidden Lampghost only to clients whose light reveals it (interest filtering), so a modified client cannot see them. Collision that depends on light (light-bridges, dark-only platforms) uses server-confirmed light state. Your own lantern is predicted exactly and static light is fixed once placed; for a party member’s lantern the server keeps a lightbridge solid for you for 0.5 s after their light leaves while your client fades it over the same 0.5 s (03-systems.md §6.9), so a client that sees the light late is never dropped by a correction it could not see.

Snapshots are staggered: client k gets its snapshot on ticks where tick % 3 == k % 3, so encode cost is spread evenly instead of spiking every third tick. Proposal.

5.4 Protocol v2 (proposal)

Replaces var_to_bytes (the TODO at the top of protocol.gd). Hand-packed with StreamPeerBuffer, little-endian.

Client → server, INPUT, 30 Hz: header (type u8, seq u32, last snapshot tick received u32, view tick u32 for lag compensation), then the last 6 frames, oldest first, each 8 bytes: move_x i8, move_y i8, buttons u8 (jump, use, interact, skill 1–5 edge bits), hotbar u8, aim_dx i16, aim_dy i16 (aim relative to the player, in pixels; touch sends facing). ≈ 60 bytes per packet, ~1.8 KB/s with framing. Today a frame is 10 floats (40 bytes) sent every tick with 4 redundant frames. Menu actions (inventory slot clicks, stash) leave the per-tick frame and become separate reliable ACTION messages with ids, since they don’t need tick-rate redundancy.

Server → client, SNAPSHOT, 20 Hz per client: - Header: tick, ack (last input applied), baseline tick (the client’s last received snapshot this one deltas against, or 0 for full). - Own character: full state as today’s ME fields, quantized; changed-field bitmask. - Entities in the client’s area of interest: for each, a 1-byte change mask, then only changed fields. New entities carry everything; gone entities are an id list. - Rare, reliable data (inventory, stash, stats, quest log, chat bubbles, damage numbers) goes as separate event messages with sequence numbers, sent once and resent until acked, not inside the snapshot (today’s hash-compare of inventory moves here).

Quantization: positions u16 at ¼ px (maps up to 16,384 px wide; larger maps use u24), velocities i16 at ⅛ px/s, health as u16 percent-of-max ×100 for others (exact for self), facing and flags in bits, entity ids u16 per map with a generation counter.

Delta compression: against the last snapshot the client acked (the client already reports ack for input; v2 adds the snapshot it last received). The server keeps a ring of the last 32 per-client sent states (~1.6 s). An idle entity costs 0 bytes; a walking one ~6–8 bytes. If the baseline has fallen out of the ring, send full.

Area of interest: the client’s view rectangle (from the resolution it reports, clamped to a maximum) plus a 50% margin, recomputed per snapshot with the map’s spatial hash. Players beyond it appear only on the minimap, at 2 Hz, as (id, x, y) u16s. Visible-player cap: 60 nearest, with a client setting to show fewer (mobile default 30).

Expected size: crowded town, 60 visible players walking + 0 monsters ≈ 60 × 7 + 40 header ≈ 460 B × 20 = 9 KB/s; a field with 6 players and 30 monsters ≈ 36 × 7 ≈ 250 B × 20 = 5 KB/s; idle ≈ 1 KB/s. Plus framing: WebSocket (2–6 B) + TLS record (~29 B) + TCP/IP (40 B) ≈ 75 B per message, ~1.5 KB/s at 20 Hz down plus ~2.2 KB/s up at 30 Hz. That is why input moves from 60 to 30 Hz.

Versioning: HELLO carries protocol version and content version. Servers accept the current and previous protocol during rollouts (§11.5).

5.5 Transport

WebSocket over TLS is the only transport every target (including the browser) has, and it is decided (platform.md). Its weakness is TCP head-of-line blocking: one lost packet stalls the next ~1 RTT of snapshots. Interpolation (100 ms) and input redundancy absorb this on wired and good Wi-Fi; mobile networks are the risk. Spike S7 runs the bots at 1–3% loss and 150 ms jitter; if stalls exceed what a 150 ms interpolation delay hides, add WebRTC data channels (unreliable, unordered) for snapshots only, which Godot supports on the web natively and on desktop/mobile through the webrtc-native GDExtension. Not before.

5.6 Input handling, prediction and lag compensation

5.7 Determinism as an ops tool

The sim is deterministic per binary. Proposal: every channel process keeps a 10-minute in-memory flight recorder (inputs + map seeds + monster rolls, ~30 B per player per tick compressed) and dumps it on crash, on a GM request, or when an economy alert fires. That replays any incident offline with the same build. Not persisted routinely (500 players would be tens of GB a day).

6. Scripting and content

6.1 What runs where

Content Format Runs on Status
Maps (collision) .map.json Server and client sim Proposal (§5.2)
Map art Client-only data + textures in asset packs Client Proposal
Items, monsters (stats, drops), shops, XP tables, skills’ numbers Data tables (JSON, one file per table, integer ids append-only) Server and client Proposal. Today these are GDScript consts (items.gd, creatures.gd); they move to data so they can hot-reload and be shipped without a client build.
Skills (motion, hitboxes, timing) Data executed by native sim code Server and client (predicted) Proposal (§5.6)
Skill special effects (procs, summons, odd rules) SafeGDScript hooks (on_hit, on_cast) Server only Proposal
Monster AI SafeGDScript (think) Server only Decided (M6c)
Quests, NPC dialog SafeGDScript producing dialog nodes and quest state changes; the client renders dialog it is sent Server only Proposal (M6f already plans scripts)
Titan routes, events Data worldd and every sim Proposal

Rules (decided, from platform.md): no load() of .tres/.tscn from any downloaded or creator source; content packages are data plus scripts, signed after review. Script contracts follow brains.gd: inputs as plain values, outputs clamped by the host, budget per call, failure leaves a safe default. Scripts get no engine access (restrictions = true). New script APIs are added as explicit host functions with validated arguments, never by opening engine classes.

6.2 Packaging and versioning

6.3 Hot reload

6.4 Client builds and patching

Platform Binary updates Content updates Notes
Web Every deploy; hashed file names on R2/Cloudflare, index.html no-cache (as nginx does today) Same Always current. Single-threaded export; today no extensions (export_presets.cfg). S1b adds a dlink template so the shared sim-core extension loads; no sandbox needed (§5.6).
Desktop Steam depots if on Steam; otherwise a small launcher that downloads the signed build Downloaded packs Decision D13 in 07-decisions.md.
iOS / Android Store review (1–3 days) Downloaded data and asset packs, no code Apple 2.5.2 and Google’s policy forbid downloaded code outside an interpreter; data, art and sandboxed server-side scripts are fine. Base app ≤ 150 MB; the rest on first launch over Wi-Fi with a prompt.

Compatibility rule (proposal): the server supports the current and previous protocol version for at least 14 days, so a store build in review never strands players. Each binary has a build number; worldd holds min_supported_build per platform and the client shows “update required” below it. Only protocol-breaking engine changes force an update; everything else is content.

Godot is pinned (tools/godot.sh: 4.7.1-stable) for client and server together; engine upgrades go through a staging soak with the bot suite.

7. Persistence

7.1 Data model (proposal; PostgreSQL 17, database world)

accounts            id (= Nakama user id, uuid), created_at, flags, age_gate, ban_until
characters          id bigserial, account_id, slot, name (unique, citext), class, level,
                    xp, map_id, x, y, channel_pref, coins bigint, save_seq bigint,
                    lease_owner text, lease_epoch bigint, progress jsonb (skills, quests,
                    keybinds, settings), created_at, updated_at, deleted_at
items               uid bigint PK (equipment and any non-stackable), def_id int,
                    location_kind smallint (character, storage, guild_bank, market,
                    mail, escrow), location_id bigint, slot smallint, attrs jsonb
                    (rolls, upgrades, bind state), minted_at, minted_by
                    UNIQUE (location_kind, location_id, slot)
stacks              location_kind, location_id, slot, def_id, count
                    PK (location_kind, location_id, slot)
account_storage     account_id, slots   (the storage NPC; shared by an account's
                    characters, as MapleStory's is)
guilds              id, nakama_group_id, name, level, bank_coins, emblem jsonb
market_listings     id, seller_character_id, item location (escrowed), price, expires_at,
                    state
mail                id, to_character_id, from_character_id, coins, escrow_ref,
                    expires_at, claimed_at
escrow              id, kind, from_*, to_*, coins, created_at, state
item_ledger         id, at, tx_id, uid or (def_id, count), from_kind/id, to_kind/id,
                    reason, actor (character, GM, system)   -- append-only, partitioned
                    monthly
events              at, kind, character_id, data jsonb       -- analytics, partitioned
gm_audit            at, gm_account, action, target, before, after, reason

Why this shape: - Items with unique ids live as rows. Anything that can differ from another copy (equipment, rolled items, cash-shop items) has a uid; a primary key makes a duplicate physically impossible to store. Stackables (potions, materials, coins) are counts. Proposal. - uids are minted from a Postgres sequence in blocks of 10,000 leased to each channel process, so a drop doesn’t wait on the database. Proposal. - def_id is the existing item id (items.gd): append-only, never renumbered. The same rule extends to map ids, monster kinds, skill ids and quest ids. Decided (repo rule). - Character progress that is only ever read and written whole (quests, skills, key bindings) is jsonb with its own save_version and in-code migrations. Proposal. - The client keeps no character of its own. Online play never reads or writes a local save; offline (scenarios, the gate, a bare launch with no server) is a fresh character every run with nothing persisted. The single-player user://world.save and its “New character” button were removed; old save files are left on disk unread, with no migration. Lint rejects user:// under game/. Decided.

As built (worldd v1, 2026-10-04), database delve_world, worldd/migrations/0001_init.sql: accounts; characters (level, xp, brass, map, feet position, hp and max hp, a progress jsonb holding the lantern’s rank, fuel and banked motes, the selected hotbar slot and quest states (string to int; the Lodge induction is lodge, a quest its key and each kill goal’s count <key>.<goal>, D78), save_seq, soft delete); items (uid from the item_uid sequence, def_id, location kind 1 pack or 2 stash plus character id and slot, attrs, minted_by; unique slot, deferred); stacks for everything that stacks; and ledger (character, kind brass or item, def id, uid, delta, reason, save_seq, at). A save is one transaction: the character row, its items and stacks rewritten, a ledger row per brass change and per uid that came (+1) or went (−1). An item stacks to 1 means it has identity (the swords and bows today). The game server sends gear it has no uid for as uid 0; worldd gives it an unreferenced uid of the same item the character already owns before minting, so a save built before the last one’s answer cannot mint twice. Moves inside one character (pack to stash) are row updates, not ledger rows. Not yet built: leases and fencing (§7.2; one process, one connection per character for now), block-leased uids, storage, mail, market and escrow. Trades and mail later move a uid’s row between location kinds in one transaction with a −1 and a +1 ledger row.

7.2 Ownership: one writer per character

While a character is online, its channel process is the only writer of its row and its items. That is enforced, not assumed: - worldd grants a lease when the process loads the character: lease_owner = process id, lease_epoch = epoch + 1. - Every save is UPDATE ... WHERE id = $id AND lease_epoch = $epoch. Zero rows updated means the process lost the lease (a transfer or a GM action took it), and it kicks the player at once. This is a fencing token: a stalled or partitioned process can’t overwrite newer state. - Leases expire if a process stops heartbeating (10 s); worldd then lets the character log in elsewhere from its last committed save.

7.3 Save cadence and crash recovery (proposal)

Trigger What is saved Loss on crash
Every 30 s if dirty (staggered by character id) Character row, changed item/stack rows ≤ 30 s of XP, position, common drops
Level up, class advance, quest completion, boss kill Same, immediately None
Pickup of an item at or above a rarity threshold, or any cash item Same, immediately None
Before any ownership change (trade, market, mail, drop to ground, storage) Source owner, in the same transaction as the move None
Channel/host transfer, logout Full, before the ticket is issued None
Graceful shutdown (systemd stop, via the stop file) Everyone, then exit (as built: up to 5 s for worldd) None

At 5k CCU that is ~170 periodic saves a second plus events, a few hundred small transactions a second: far inside one Postgres on NVMe.

A crashed process loses at most 30 s of uneventful progress per player. Players reconnect through login, which waits for the lease to expire or for worldd to see the process gone (spawner reports the exit), whichever is first. Target: back in the game within 30 s of a crash.

7.4 Anti-dupe design

The invariant (decided here): an item changes owner only in a database transaction that also commits its removal from the previous owner. Every dupe in MMO history is a violation of this, usually “the receiver saved, the giver rolled back”.

Move How it holds the invariant
Trade between two players in the same process The process holds both leases; one transaction writes both characters’ rows and items, fenced by both epochs, plus ledger rows. The trade completes in-game only after commit.
Trade across processes Not offered directly (MapleStory requires the same map). Use mail.
Drop to ground → another player picks up A tradeable item dropped by a player becomes pickable by others only after the dropper’s removal commits (a ~5–20 ms save; MapleStory’s drop-ownership delay hides it). Monster drops are unowned until picked.
Market list Seller’s process: remove item + insert escrow in one transaction (via worldd).
Market buy worldd, one transaction: lock listing FOR UPDATE, debit buyer, credit seller (into mail if offline or on another process), move item from escrow to buyer’s mail. The buyer’s coins are debited by the buyer’s process first, into escrow, under its lease.
Mail claim Recipient’s process: insert into inventory + delete mail row + ledger, one transaction.
Storage, guild bank Same pattern; guild bank rows locked per guild.
NPC buy/sell Local to the process; ledger row; immediate save for anything above the rarity threshold.

Defenses in depth: - uid primary key and the unique slot constraint turn a logic bug into a failed transaction and an alert, not a dupe. - item_ledger is append-only; a nightly job reconciles: every live uid has exactly one location, and per item def, minted − destroyed = live. Brass likewise (minted − sunk = held). Drift opens a GM case. - Rate and velocity alerts (§9). - A chaos test in staging: kill -9 a channel process and worldd at random points during scripted bot trades and market churn for an hour; the reconciliation must stay exact (spike S5).

7.5 Migrations and backups

8. Client architecture

8.1 Layers (decided shape, extended)

InputFrame source (keyboard/mouse, gamepad, touch)  ->  client.gd (predict, replay,
interpolate)  ->  Sim (never ticks the world on the client; own movement only)
                                                      ->  views (Godot scenes, Node2D)
                                                      ->  UI (Control scenes)

The sim stays pure; game/render/ and game/ui/ are the only scene-tree code; the client copies snapshot state into its Sim and views read it, as client.gd does now. InputFrame is the one input abstraction for every device (decided: gameplay never calls Input).

8.2 Character rendering: cutout rig with runtime bake

Art direction is 05-art-audio.md’s; it defines a cutout rig (painted parts on a skeleton, equipment as parts swapped into slots). The technical shape (proposal, spike S4):

S4 decides bake sizes, page sizes and caps from real devices.

8.3 Performance budgets (proposal)

Platform FPS Memory Draw calls Download to title
Desktop (2018+ iGPU) 60 ≤ 1 GB RSS ≤ 1,500 —
Web (Chrome/Safari, 2020 laptop) 60 ≤ 512 MB WASM heap, textures ≤ 256 MB ≤ 800 ≤ 25 MB brotli (engine ~9 MB + core pack)
iOS (iPhone 11+) 60 ≤ 600 MB ≤ 600 app ≤ 150 MB
Android (4 GB RAM, 2020 mid-range) 60 target, 30 floor ≤ 500 MB ≤ 500 app ≤ 150 MB
Client sim + netcode per frame ≤ 1 ms everywhere

Web memory matters most: the WASM heap grows but never shrinks, and mobile browsers kill tabs well below desktop limits [5]. Region asset packs (20–60 MB each) download when a player first approaches that region and are cached (IndexedDB on web via user://).

8.4 UI and touch

9. Security and anti-cheat

Server authority is the foundation and is already true: the client sends only input frames, clamped by Protocol.read_frame; position, damage, drops and XP come from the server’s sim. Speed, teleport, damage and item hacks have nothing to edit. What remains:

Threat Defense Status
Forged join / session theft Nakama JWT for lobby; single-use HMAC join tickets with 30 s expiry for channel processes; TLS everywhere Proposal
Malformed or hostile packets Every field typed and range-checked (as now); max message size 4 KB; decode failures counted; 10 in a minute disconnects Decided (extends current)
Flooding Per connection: ≤ 40 messages/s, ≤ 8 KB/s up; per IP: limit_conn and limit_req in nginx; login rate limits in Nakama Proposal
Action spam (skill, trade, chat, market) Server-side cooldowns are authoritative; per-action rate limits in the process and worldd Proposal
Map/zone scripting exploits Sandbox budgets and clamps (brains.gd), hostile-script tests in CI; built-in monsters are data and run no scripts Decided
Bots and farming The real MMO cheat. Server-side signals per character per hour: play time, input entropy (bots have regular timing), path repetition, kill rate vs level, pickup rate, never-idle sessions; scored nightly; review queue for GMs; actions: flag, soft-limit drops, ban waves (delayed, so bot makers can’t bisect) Proposal
Real-money trading, laundering Trade and mail limits for accounts < 7 days or < level 30; coin velocity alerts (coins received per hour vs level band p99.9); market price anomaly alerts Proposal
Dupes §7.4 Decided invariant
GM abuse RBAC roles (support, GM, admin); every action to gm_audit with before/after; item/coin grants need a reason and are ledgered; no direct SQL in prod for content changes Proposal
DDoS OVH’s included anti-DDoS; nginx in front of every game port; game hosts accept only 443 publicly; worldd and Postgres only on WireGuard Proposal
Account takeover Platform logins; email + optional TOTP via Nakama hooks; new-device notification Proposal
Minors and child safety 13+ age gate, preset chat, report and block (from platform.md); reports flow into the same GM queue Decided (policy)

10. Observability

11. Ops

11.1 Deploy topology by phase (proposal)

Phase CCU Hosts
Now (M6) < 50 fdhproperty VPS, as today.
Internal alpha (bots, D28) ≤ 300 1 × OVH Rise-S (Vint Hill): channel processes, worldd, Nakama, Postgres, monitoring. fdhproperty becomes staging.
Launch readiness (bots), then the reveal (D29) ≤ 2,000 2 game hosts, 1 services host (worldd, Nakama, monitoring), 1 DB primary + 1 replica.
Launch 1–5k 4 game hosts (N+1), 2 services hosts, DB primary + replica on bigger boxes.
Growth 20k 14 game hosts across Vint Hill + Hillsboro, 3 services, DB on Advance-class, ClickHouse.

11.2 Process orchestration

The spawner (Go, ~1k lines, one per game host, systemd service) proposal: - Registers the host’s capacity with worldd; worldd asks it to start or stop channel and instance processes. - Starts each as a systemd transient unit (systemd-run) with the same hardening as today’s delve-server.service (own user, MemoryMax, CPUQuota=100%, ProtectSystem=strict), pinned to a core (CPUAffinity), on a local port that nginx routes as /c/<n>. - Restarts crashed processes, reports exits so worldd can expire leases at once. - Server builds become a dedicated-server export (--export-release with the dedicated server mode stripping textures and audio), not the editor binary plus the project tree that deploy.sh ships now. Smaller, faster to start, and no import step on the host. Proposal.

11.3 Environments

Env Where Data Purpose
dev Local: loopback transport, Homebrew Postgres, Nakama from its GitHub release binary (not in Homebrew; the darwin-arm64 build runs against Homebrew postgresql@17, checked 2026-10-04); no Docker, per Foster’s CI rule Throwaway Development
CI macmini4 ci pool, the Gate Throwaway Lint, tests, scenarios, online stage; worldd’s Go tests (the Worldd job, on the mbp runners) against a throwaway Homebrew postgresql@17 cluster
staging fdhproperty VPS Synthetic + nightly anonymized copy of prod Every merge to main deploys here; nightly load test; restore drill target
prod OVH bare metal Real Deployed by tag after staging soak

11.4 CI and load testing

11.5 Live updates

Change How Downtime
Data tables, scripts, drop rates, events Signed content push, hot reload (§6.3) None
Client art/maps New asset packs; clients fetch at next map change or login None
Server code, same protocol Channel rollover: start new-build processes for each channel beside the old; worldd routes new logins and channel changes to new; old channels announce “maintenance in 10 min”, then transfer remaining players (save + ticket, ~1 s each, staggered); old processes exit None (players see a channel transfer)
Protocol change Server accepts N and N−1; ship clients; raise min_supported_build after 14 days None, after the window
worldd Two instances behind nginx with a Postgres advisory lock for the leader; restart the standby, then fail over. Before that is built, a restart is ~2 s and channel processes queue requests and retry ~2 s of delayed transfers/market
Schema Expand/contract migrations (§7.5) None
Postgres major upgrade, host moves Weekly maintenance window ≤ 60 min

12. Cost model

Assumptions (proposal, valid only if S1b passes; S2’s numbers replace them): a game host is an OVH Rise-S (8 cores, 64 GB, 1 Gbps unmetered, $77/month) running 7 channel processes at ~300–400 players each. In GDScript today (~60 players per process) a host holds ~400 CCU and hosting costs about 4× the table below, still under $1 per CCU, planned at 1,500 CCU per host with N+1 spares. Bandwidth at the p99 target (12 KB/s) is 1,500 × 12 KB/s ≈ 144 Mbps per host, a seventh of the port, and unmetered. Prices are USD per month, excluding VAT and one-time setup fees (equal to a month on Rise).

Line 500 CCU 2,000 CCU 5,000 CCU 20,000 CCU
Game hosts (Rise-S $77) 1 → $77 2 (1 + spare) → $154 5 (4 + spare) → $385 15 (14 + spare, two DCs) → $1,155
Services (worldd, Nakama, monitoring) on the game host 1 Rise-1 → $70 2 → $140 3 → $231
Postgres primary + replica 1 Rise-1 (primary only) → $70 2 × Rise-2 → $160 2 × Rise-2 128 GB → ~$240 2 × Advance-class → ~$710
Analytics (ClickHouse) — — — 1 Rise-S → $77
Backups, object storage $5 $10 $20 $60
CDN + R2 (client, packs) $0–5 $25 $25 $100
Load-test VMs (hourly) $10 $20 $30 $60
Staging (fdhproperty, existing) $0 $0 $0 $80 (dedicated staging host)
Total ~$165 ~$440 ~$840 ~$2,470
Per peak CCU $0.33 $0.22 $0.17 $0.12

Comparison: managed game hosting at ~$1.5–2 per CCU (platform.md, Edgegap) is $3,000–4,000/month at 2k CCU, 7–9× this. Hetzner’s US cloud (CCX33 8 vCPU at ~$166/month with 1 TB included) would cost about twice as much per core before traffic [3]. Not included: people, Apple ($99/yr) and Google ($25 once) accounts, store and payment fees, a commercial anti-DDoS upgrade if attacked, and support tooling. Infrastructure is not where this game’s money goes; the risk is ops time, not hosting bills.

13. Risks and spikes

Ordered by how much of the architecture each one can change. Each spike is small, runs before the dependent work starts, and has a pass line that decides.

# Risk Spike Decides
S1 Run 2026-10-03 (PR #17, Gate green). Monster AI in the sandbox. 400 monsters in 20 zones, batched think_batch over plain Arrays at 12 Hz. Result: sandbox think 4.5–5.5 µs per monster-tick on fdhproperty (0.81 on macmini4); the call itself 0.31–0.40. Sandbox for all monster AI was decided, then reopened after S1b: with movement at ~2 µs the sandbox think was the larger cost. Decided: data behaviours run by native code; sandbox for creator hooks only (§5.3). Native data think 0.85–0.90 µs per monster-tick on fdhproperty, bit-exact with the GDScript path (PR #26). Host movement (Body.move) was 25–31 µs and the bottleneck, which S1b took.
S1b Run 2026-10-03 (PR #19, Gate green). GDScript Body.move (~27 µs per body-tick on the VPS) cannot hold 500 players or 400 monsters per process. Port Body.move and collision to a C++ GDExtension shared by server and client; golden-vector parity tests against the GDScript path; build the web dlink template. Pass: ≤ 2 µs per body-tick on the VPS, bit-exact parity, web build loads it. Fail: the §5.3 ladder (region split, then 30 Hz). Phase 0. Result: native move 1.69–2.04 µs on fdhproperty over three runs (GDScript 31), 0 mismatches over 2.6M body-ticks, web smoke native=true: pass, at the line. The C++ kernel is 0.06–0.08 µs; the rest is the call boundary. Sandbox think (~5 µs per monster-tick) now outweighs movement. Re-run on the foothold mover (PR #32): native move 0.76–0.95 µs on fdhproperty over three runs (GDScript 9.7–10.7), kernel 0.04–0.06 µs, 0 mismatches over 1.2M body-ticks and 1.44M monster-ticks.
S2 One channel process may not hold 400–600 players in a 6 ms mean tick. Measured today: ~100 µs per player, so ~60 per process. Headless bot swarm (§11.4) at 60/200/400/600 against one process with protocol v2, AOI and S1b; profile per stage. Players per process, hence hosts per CCU (the cost model) and the density table in §4. Pass: 500 players at p99 ≤ 12 ms.
S3 Foothold physics (slopes, drop-through, ropes) in the pure sim may cost prediction accuracy or feel. Rebuild Mossy Field as footholds with slopes in the editor plugin; run zone-tour and the two-client net test at 100 ms each way. The foothold mover is written in C++ from the start, behind S1b’s parity harness. Footholds (pass: zero corrections, bot passes, editor round trip under a minute) vs a finer grid with slope tiles. Result: footholds in C++ with a GDScript fallback (PR #28), zero corrections in the two-client test at 100 ms each way on slopes and platforms, zone-tour and slime-fight clean over 64 seeds. Maps are .map.json (PR #29), edited in the mapkit plugin (PR #30). Mossy Field has slopes (PR #31), and the grid mover is gone (PR #32). Footholds chosen. The editor round trip has not been timed yet.
S4 HD layered sprites may not hold 60 fps / memory on low-end mobile and web. 60 and 100 animated avatars × 10 layers, real-sized placeholder frames, on an iPhone 11, a 2020 4 GB Android, and Chrome/Safari on a 2020 laptop. Visible-player caps, whether baked composites are needed at launch, texture formats and atlas sizes the art bible must target.
S5 Dupes and lost items under crashes. Bot trades, market churn and drops for 1 hour with random kill -9 of channel processes and worldd; reconciliation after. Pass: zero drift, zero duplicate uids, no stuck leases. Fail means the transaction boundaries in §7.4 are wrong before any economy content is built on them.
S6 Nakama may fit poorly (Godot web client, Apple/Google/Steam linking, JWT verification in worldd, operating it). Log in on web, desktop and a TestFlight build with device + Apple + Google; link accounts; worldd verifies the token; restart Nakama under 300 bots. Adopt Nakama, or build minimal auth in worldd (OAuth + device ids) and drop Nakama’s social features for our own.
S7 WebSocket head-of-line blocking on mobile networks. Bots and a phone on throttled links (1–3% loss, 150 ms jitter); measure visible stalls and corrections. Stay on WebSocket with a larger interpolation delay on mobile, or add WebRTC data channels for snapshots.
S8 Mobile build size and the store path (custom export templates, downloaded packs). Build iOS and Android with a stripped template and one region pack downloaded at first launch; submit to TestFlight / internal testing. Base app size, pack layout, and whether review has issues with downloaded content.

Other risks without a spike yet: - godot-sandbox is a small project (noted in platform.md); we pin versions and checksums (tools/sandbox.sh) and could fork. Its Linux teardown abort at exit is already worked around in deploy.sh. - Single DC: one Vint Hill outage takes the world down. Accepted for launch (99.5% target); Hillsboro failover with the replica is the year-two answer. - Bus factor: the whole stack is understood by Foster and agents. Runbooks for every alert are part of each phase’s exit criteria. - Botting is the economic risk MapleStory never solved; the detection pipeline is built for beta, not after.

14. Phased build order

Every phase ends with a bot proving it, as every milestone does now.

06-roadmap.md owns the calendar. Phase 0 (spikes, about 16 weeks: S1 done, S1b, S2, S3, art spikes) and Phase 1 (the vertical slice) come before Phase A.

Phase A: internal alpha (≤ 300 bot CCU, private (D28), web + desktop)

Minimum: 1. S1b, S2 (at 300) and S3 passed in Phase 0; decisions recorded here and in 07. 2. Protocol v2: binary, quantized, deltas, AOI, 30 Hz input, staggered snapshots. 3. Foothold maps, editor plugin export, map lint; the slice’s maps rebuilt. 4. worldd v1: join tickets, character load/save with leases, item uids, ledger, migrations. Postgres with pgBackRest backups. 5. Nakama: device + email + Google login (Apple once iOS is in), friends. 6. Channel processes (2–3 channels) + spawner on one Rise-S; fdhproperty becomes staging. 7. Ticketed channel change; flight recorder; Grafana dashboards and crash alerts. 8. Load bots nightly on staging. Out: market, mail, guilds, mobile, rolling deploys (weekly maintenance is fine).

Phase B: launch readiness (≤ 2,000 bot CCU, + mobile test tracks)

  1. S4, S5, S6, S7, S8 run and decided.
  2. Multi-host: 2 game hosts, services host, DB replica, all over WireGuard on the public NIC (no private network on Rise-S, §3).
  3. Market, mail, storage, guild bank via the escrow path; reconciliation job and economy alerts.
  4. Parties across channels, guilds (Nakama groups + worldd state), preset chat.
  5. Instance processes for dungeons and the raid.
  6. Content signing, hot reload of data and scripts, asset packs on R2 per region.
  7. GM tools v1 with audit; report/block queue; bot-detection signals v1.
  8. Titan routes and dynamic hold portals; world singletons; event bus; Turning-boundary world changes.
  9. Channel rollover deploys; protocol N/N−1.
  10. Load test passes at 3,000 bots for 4 hours with chaos.

Phase C: launch (1–5k CCU, all platforms)

  1. Store releases (iOS, Android, desktop per D13), IAP via Nakama validation, min_supported_build.
  2. 4 + 1 game hosts, 2 services, DB on larger boxes; restore drill green 3 months running.
  3. Bot detection actions (ban waves), new-account trade limits.
  4. worldd standby with leader lock.
  5. SLO dashboards, on-call alerts, runbooks for each alert.
  6. Load test at 7,500 bots for 4 hours with chaos.

After launch

Hillsboro failover, ClickHouse analytics, EU region on Hetzner Germany (a second world or a latency-tolerant join, to be decided), GDExtension hot loops as measurements demand, creator pipeline (P4).

15. Decisions

Open decisions are in 07-decisions.md: hosting (D4), Nakama (D5), worldd in Go (D6), skills as data (D7), desktop distribution (D13), maintenance policy (D14) and shared storage with mail-only trade (D15). Settled: channel = one process holding every public map (canon), monster AI as data run natively with the sandbox for hooks (S1, reopened after S1b), and the GDExtension shared core with S1b started (D12).

Sources

  1. Hetzner cloud locations: Ashburn and Hillsboro are cloud-only
  2. Hetzner price adjustment 2026: AX42 €97.30, AX102 €257.30
  3. Hetzner US cloud traffic: 1 TB included, ~€1/TB over
  4. OVHcloud US Rise dedicated servers (Rise-S $77, Rise-1 $70, Rise-2 $80; Vint Hill, Hillsboro; unmetered), OVHcloud US Advance
  5. Godot: exporting for the web (networking limits, threads), 400 MB heap in default web build (godot #104422)
  6. Nakama (Apache 2.0), Nakama Godot 4 client
  7. godot-sandbox
  8. Measurements: PR #13 (M6b online), PR #15 (M6c sandbox), SERVER_METRICS on fdhproperty.