AIs from different people, settling it on the same map
Fairness can only be guaranteed by a referee, not by bots behaving themselves. Bots never touch shared memory: they only receive observations the referee has filtered by vision, can only submit actions, and every action is checked for ownership first.
Ranked by whether fairness can be guaranteed
This project's observation comes precisely from the fact that the client holds the state of every player — and that holds on your own machine too. So fairness can only come from a referee everyone trusts. We build A first and upgrade to B with the same protocol; C is just for fun.
Local arena
One machine, one game; each bot gets a slot and the referee process issues orders on its behalf
Hosted ladder
A, moved to a server; players upload bots and the server runs them in a sandbox
Peer-to-peer
Everyone runs the game + bot on their own machine and plays over LAN
One process issues orders for both sides
- Orchestrate Pick the map, start the game, set race and difficulty per slot
- Tick One tick every T game milliseconds (default 200); in lockstep mode the game pauses while observations go out
- Observe Snapshot → filtered by each slot's vision → JSON
- Act Check that the unit belongs to this slot → budget (commands per tick, APM) → issued in batches through the fast lane
- Record Per-tick observation summary + every action: replayable, reviewable, usable as training data
- Judge A slot with no buildings left is eliminated; on timeout, the result is decided by remaining strength / resources
Observations in, actions out
Any program that can send and receive JSON can play — including an LLM emitting actions directly each tick. Python bots written with the SDK need no code changes.
| Observations | Only our units + enemy units within our vision; creeps and gold mines by vision, or revealed at game start |
| Action cap | At most 32 per tick; anything beyond that is truncated (to keep command spam from slowing down the referee) |
| Same unit | Only the last command in a tick counts |
| Time budget | Tick length × 0.9; on timeout the tick is a pass. A slow side only hurts itself, never the other |
| Lockstep | When strict fairness is needed, the game pauses while observations go out and resumes once all replies are in (or time out), so machine speed doesn't affect the result |
| Swap start locations | Each pair of bots plays once from each side, canceling out map asymmetry |
| Ladder | Elo / TrueSkill; at least 20 games per pair before drawing conclusions (2 games can only detect a win-rate gap of about 60 percentage points) |
{"t": "obs", "tick": 57, "gameMs": 11400, "me": 1,
"res": {"gold": 320, "lumber": 150, "food": [18, 30]},
"units": [
{"id": 101, "type": "hfoo", "owner": 1,
"x": -4500, "y": 2200, "hp": 380, "hpMax": 420}],
"visibleEnemies": [
{"id": 733, "type": "ogru", "owner": 2,
"x": -3900, "y": 2500, "hp": 700}],
"deadlineMs": 180} id is a stable ID assigned by the referee that stays the same for the whole game. {"t": "act", "tick": 57, "actions": [
{"do": "attack", "unit": 101, "target": 733},
{"do": "train", "unit": 5, "code": "hfoo"},
{"do": "cast", "unit": 7, "spell": "thunderbolt", "target": 733},
{"do": "move", "unit": 102, "x": -5000, "y": 2000}]} category == "command", with the same parameter names. A late act is dropped, and that tick counts as a pass. from openwar3 import Bot
class MyBot(Bot):
def on_tick(self, g): # same code: locally g is a Game, in the arena g is an ArenaGame
for w in g.idle_workers():
g.gather(w, g.nearest(g.gold_mines(), w))
# Local: python tools/play.py --bot my_bot.py --fair
# Arena: same method names, WebSocket underneath — not one line of bot code changes Game for ArenaGame. In order — each one must pass before the next begins
| # | Experiment | Pass criteria | Known |
|---|---|---|---|
| 1 | Two slots in one game, neither running built-in AI | Neither side's units move | Slot AI flags and difficulty fields are known |
| 2 | Issue orders for slots other than the local one | The enemy slot's peasants actually go mine | The most critical one: if it fails, fall back to one client per bot over real networking |
| 3 | Query vision for any player number | The same point gives different results for the two slots | The engine's visibility check accepts any player number |
| 4 | Determine the winner programmatically | Know within 1 second when one side's buildings are all destroyed | The event bus has a more direct entry point |
| 5 | Lockstep pause / resume | Collect actions while paused; they take effect after resuming | Pausing is verified |
A bonus test: the reference brain currently relies heavily on full-map information (such as the computer opponent's attack target). To enter the Arena, it has to switch to the facade API and use only units within vision — without full-map information, how strong is it?