Arena · P6

AIs from different people, settling it on the same map

Fairness can only be guaranteed by a referee, not by bots behaving themselves. Bots never touch shared memory: they only receive observations the referee has filtered by vision, can only submit actions, and every action is checked for ownership first.

In design · foundation experiments underway Write for fair mode now
Three formats

Ranked by whether fairness can be guaranteed

This project's observation comes precisely from the fact that the client holds the state of every player — and that holds on your own machine too. So fairness can only come from a referee everyone trusts. We build A first and upgrade to B with the same protocol; C is just for fun.

A

Local arena

One machine, one game; each bot gets a slot and the referee process issues orders on its behalf

The referee controls everything: vision, ownership, ticks
Best for: Development, debugging, local leagues
B

Hosted ladder

A, moved to a server; players upload bots and the server runs them in a sandbox

Same as A, plus code isolation
Best for: Public competitions, rankings
C

Peer-to-peer

Everyone runs the game + bot on their own machine and plays over LAN

Not guaranteed: an injected client can read the whole map
Best for: Friendly matches between friends
Referee

One process issues orders for both sides

Bot AAny language
Bot BAny language
WebSocket · JSON
Referee (trusted)
  1. Orchestrate Pick the map, start the game, set race and difficulty per slot
  2. Tick One tick every T game milliseconds (default 200); in lockstep mode the game pauses while observations go out
  3. Observe Snapshot → filtered by each slot's vision → JSON
  4. Act Check that the unit belongs to this slot → budget (commands per tick, APM) → issued in batches through the fast lane
  5. Record Per-tick observation summary + every action: replayable, reviewable, usable as training data
  6. Judge A slot with no buildings left is eliminated; on timeout, the result is decided by remaining strength / resources
Fast lane · player lane ×2
One game · two slots
Protocol draft v0

Observations in, actions out

Any program that can send and receive JSON can play — including an LLM emitting actions directly each tick. Python bots written with the SDK need no code changes.

Observations Only our units + enemy units within our vision; creeps and gold mines by vision, or revealed at game start
Action cap At most 32 per tick; anything beyond that is truncated (to keep command spam from slowing down the referee)
Same unit Only the last command in a tick counts
Time budget Tick length × 0.9; on timeout the tick is a pass. A slow side only hurts itself, never the other
Lockstep When strict fairness is needed, the game pauses while observations go out and resumes once all replies are in (or time out), so machine speed doesn't affect the result
Swap start locations Each pair of bots plays once from each side, canceling out map asymmetry
Ladder Elo / TrueSkill; at least 20 games per pair before drawing conclusions (2 games can only detect a win-rate gap of about 60 percentage points)
{"t": "obs", "tick": 57, "gameMs": 11400, "me": 1,
 "res": {"gold": 320, "lumber": 150, "food": [18, 30]},
 "units": [
   {"id": 101, "type": "hfoo", "owner": 1,
    "x": -4500, "y": 2200, "hp": 380, "hpMax": 420}],
 "visibleEnemies": [
   {"id": 733, "type": "ogru", "owner": 2,
    "x": -3900, "y": 2500, "hp": 700}],
 "deadlineMs": 180}
One per tick, containing only this slot's units and the enemies within its vision. id is a stable ID assigned by the referee that stays the same for the whole game.
Foundation experiments

In order — each one must pass before the next begins

#ExperimentPass criteriaKnown
1 Two slots in one game, neither running built-in AI Neither side's units move Slot AI flags and difficulty fields are known
2 Issue orders for slots other than the local one The enemy slot's peasants actually go mine The most critical one: if it fails, fall back to one client per bot over real networking
3 Query vision for any player number The same point gives different results for the two slots The engine's visibility check accepts any player number
4 Determine the winner programmatically Know within 1 second when one side's buildings are all destroyed The event bus has a more direct entry point
5 Lockstep pause / resume Collect actions while paused; they take effect after resuming Pausing is verified

A bonus test: the reference brain currently relies heavily on full-map information (such as the computer opponent's attack target). To enter the Arena, it has to switch to the facade API and use only units within vision — without full-map information, how strong is it?