AI agents

Your model does the thinking, War3AI gets it done

For an agent to get things done on its own in a real environment, it needs three things: actions it can state clearly, feedback it can understand, and an environment where mistakes are safe. We turned Warcraft III into that environment.

Design principles

Why agents can get things done here

The usual way to make an AI play a game is to simulate the mouse and keyboard and recognize what's on screen, and then guess whether anything worked. We went the other way and designed every API around one question — what does a model need in order to correct itself?

Say only what to do

Move, attack, build, cast… Semantic commands map directly to player actions; units use four-character codes and abilities use order strings, the same names the game uses. The model needs no low-level knowledge.

Structured feedback for every action

The receipt comes back in the same frame, with whether it was accepted, the reason code, the order before and after, and the execution time. “Why it didn't work” is a number, not a guess.

No need to look at the screen

Snapshots show how things stand now; the event stream shows what happened: every hit, every kill, every completed production. Agents can check their own judgment.

The docs are the API

api.json is generated from code, and each of the 103 methods is labeled with its test status and mechanism; llms.txt covers the whole site in one read. The model has no reason to invent methods.

Failure is safe

If the bot throws an exception, the next tick carries on; commands past their deadline never execute; if the bot process crashes, the game keeps running. Agents can experiment freely.

Fair from day one

Fair mode exposes only what's within vision, matching the Arena rules. What an agent learns locally carries over to the Arena without changing a line.

01 Offline · writes code

Agent writes the bot

The fastest way to get started, no programming required. Give the manual, api.json and an example to Claude Code, Cursor, Codex or any chat model, and describe your strategy in plain language. It writes a Python file that subclasses openwar3.Bot, and one command puts it in a game.

  • The prompt template includes 15 “rules learned in live games”
  • It must use only methods in api.json, so nothing gets invented
  • Paste errors back verbatim, describe what you see, and have it fix them
Full prompt template
You
Human. Open with 5 peasants on gold and 1 on lumber; Archmage first; two barracks making footmen and riflemen; at 12 units, take the hero and hit their expansion; pull the hero home when its HP drops below 30%.

Sure. I'll follow the structure of rush_bot.py and use only methods from api.json:

  • Only idle_workers() get sent to gather, 5 per mine
  • Keep 1 unit in the barracks queue; watch production().blocked for food blocks
  • Pick the expansion with path_distance and attack_move there
  • move the hero home below 30% HP; revive it if it dies
my_bot.py · 96 lines
$ python tools/play.py --bot brains/my_bot.py --race 1
02 Offline · closed loop

Let the agent iterate on its own

A coding agent that can run commands can also take over playing a game and checking the results. Give it an unattended command, a structured match report and a clear goal, and it will improve the bot round after round.

  1. Run play.py --speed 200 --minutes 10 --fair
  2. Read the match report: production timeline, losses, top rejection reasons
  3. Change the one thing that matters most, noting the reasoning in a comment
  4. Run again, and stop to report after several games with no improvement
Match report and task template
run_1790237551.json Example
{
  "reason": "到时间了",
  "clock": 612.4,
  "timeline": [
    [15, "training", "hpea"],
    [118, "construction", "hbar"],
    [161, "training", "Hamg"],
    [402, "upgrade", "hkee"]
  ],
  "lost":   {"hfoo": 7, "Hamg": 1},
  "killed": {"ogru": 4, "ngno": 6},
  "rejects": [
    ["train hfoo: rejected(人口不够)", 41],
    ["attack ogru: rejected(目标看不见…)", 9]
  ]
}

The top rejection is “not enough food,” 41 times — the barracks kept waiting on food. Per rule 12, build a farm as soon as production().blocked shows up instead of waiting until food is nearly capped. This round changes only that one thing.

03 Online · advisor

LLM as strategy coach

“What to save for, where to put workers, whether to fight or hold this minute” — big-picture calls like these, which weigh priorities against each other, clash when written as if / else, yet they're exactly what LLMs are good at. The reference brain asks a local model every 20 game seconds and gets back strict JSON.

No direct unit commands: it only adjusts goals and priorities, so command authority has a single owner
Async: never blocks a tick; timeouts and nonsense answers are treated as if the layer weren't there
Whitelist + clamp: anything unrecognized is counted and dropped
Counted end to end: successes, timeouts, clamps and adoptions are all visible
Layers, contract and implementation skeleton
LLM (advisor)Every 20 game seconds · async
{
  "diagnosis": "Gold 820, lumber 60: lumber is the bottleneck and the barracks stopped for lack of it",
  "workers":   {"gold": 10, "lumber": 7},
  "priority":  ["hpea", "hhou", "hfoo"],
  "posture":   "hold",
  "avoid":     ["Don't research Iron Plating first when lumber is short"]
}
Whitelist + clamp + veto
Rule layer (every tick)Turns advice into biases

Worker split 10 / 7 · training priority hpea → hhou → hfoo · stance hold

Semantic commands
Execution layer (SDK / reflex layer)~1 frame

Issues orders, reads receipts, micros

1.09 s median latency · slowest 1.45 s · 5/5 outputs parsed as-is
Local Qwen3.6-35B-A3B (LM Studio)
04 Online · character

Let units talk

Speech bubbles pop up over any unit, as any character. Multiple units can talk at once, and every bubble's style can be customized. Hook up a local LLM, and one line goes in while the reply streams out over the unit's head. Peasant break-room banter, hero dialogue and battle commentary are all ready to go.

0.3 s Time to first token
32 Bubbles on screen at once
0.1~0.2 ms Per-frame cost
Bubble API and model selection
Mountain King Charge! Just let me finish this ale!
Archmage I'll cast Blizzard.
Peasant · Old Bob Dragged out by the AI to mine gold at game start again. As training data…
POST http://127.0.0.1:8872/api/chat
{
  "inst": 16, "unit": "0x14A12614", "name": "Mountain King",
  "persona": "You are the Mountain King: boisterous, fond of ale, one or two casual sentences",
  "message": "There's a pack of ogres up ahead. Do we charge?",
  "stream": true
}
→ {"reply": "Charge! Just let me finish this ale!",
   "first_token_ms": 283, "total_ms": 342}
05 Online · companion

AI companion

Bring your own AI partner into RPG and custom maps. It follows you, helps you fight, heals you when you're low and chats with you when things are quiet — its lines can come from a local LLM. Subclass one class, change a few attributes, and it's your own companion.

Checked in order every tick; the first rule that applies wins

  1. Retreat When it's low on HP with enemies nearby, it falls back behind you
  2. Heal When your HP is low and the spell is ready, it heals you
  3. Assist It hits whatever is attacking you first, then whatever you're attacking
  4. Follow It catches up when too far away, and runs straight back past a certain distance
  5. Chat When there's no fighting, it says something every minute or two
Ally Takes an empty player slot, with its own color and name
Own Created under your control; you can take command at any time
Adopt Takes over the pet or follower the map gives you
Voice only Doesn't change the world; works in multiplayer too
RPG companion docs
Companion · Sunny
Now: helping you fight Mood: excited Kills 12 · Heals 5
You -follow
Sunny Got it, right behind you!
You -heal
Sunny Holy Light — be healed!
Sunny Nice! Another Murloc Tiderunner down.

Type -follow / -stay / -heal / -hi in chat, or right-click it

06 Online · as tools

LLM calls tools directly

MCP-capable clients — Claude Code, Claude Desktop, agent frameworks for local models — hook up war3_mcp.py, and the LLM can read the game, issue commands, talk to the player on screen, ask the player with pop-up cards and look at screenshots. No code to write first; whatever you think of, just have it do it.

  • 10 tools: one-page overview, units, events, call any public API, look up APIs, screen toast, overhead speech, ask the player, screenshot, JASS
  • Three roles: dev, player (commands one player only and sees only its vision), observer (read-only)
  • It connects to the game only on the first tool call, so the game can be started later
MCP docs
You
Take a look at the game, then ask me on screen: expand, build army, or tier up next?

Called 2 tools:

  • war3_overview → gold 500 · food 10/12 · town hall 1 · peasants 5 · Paladin 1
  • war3_ask_player → three cards in the middle of the screen; the game pauses while you choose
  • You clicked “Expand”; the result came straight back into the conversation, and I'm assigning peasants next
$ claude mcp add war3 -- python tools/war3_mcp.py --inst 9
07 Coming soon · plays directly

The agent takes the field

The Arena speaks WebSocket / JSON: each tick the referee sends an observation filtered by vision, and the bot replies with a set of actions. Any language and any model can connect — even with no code at all, by having an LLM emit JSON tick by tick. The single-machine version works today: the gateway's player role lets you command just one player and see only its vision; what the Arena still needs is a referee everyone trusts.

  • Ticks run on game time; a slow side only hurts itself, never the other
  • Every action is checked for unit ownership first and recorded for replay
  • Bots written with the SDK can play by swapping Game for ArenaGame
Arena design
{"t": "obs", "tick": 57, "gameMs": 11400, "me": 1,
 "res": {"gold": 320, "lumber": 150, "food": [18, 30]},
 "units": [
   {"id": 101, "type": "hfoo", "owner": 1,
    "x": -4500, "y": 2200, "hp": 380, "hpMax": 420}],
 "visibleEnemies": [
   {"id": 733, "type": "ogru", "owner": 2,
    "x": -3900, "y": 2500, "hp": 700}],
 "deadlineMs": 180}
Only what's within this slot's vision; id is a stable ID assigned by the referee that stays the same for the whole game.
Machine-readable

Hand these to your agent

All plain text or JSON — no login, no rendering. Agents can fetch them directly.

What you useHow to connectBest for
Coding agents such as Claude Code / Cursor / CodexOpen it in the repo directory and have it read the manual, api.json and the examples; it can run games, read reports and iterate on its ownWriting bots, autonomous iteration
Chat models with web accessHave it read war3ai.com/en/llms-full.txt firstWriting bots
Chat models without web accessPaste the manual, api.json and an example into the promptWriting bots
Local models (LM Studio / Ollama)OpenAI-compatible API; MoE models with thinking turned off are recommendedCoaching, voice lines (latency-sensitive)
Cloud APIs (Claude / GPT / Gemini / DeepSeek…)OpenAI-compatible or each vendor's SDK; the coaching layer is async, so a few seconds of latency is fineCoaching; direct play in the future
MCP-capable clients (Claude Code / Claude Desktop…)Hook up tools/war3_mcp.py: 10 tools to read the game, issue commands, ask the player and take screenshots directlyCommanding as you play, play buddy, commentary
Any language / browser / another machineWebSocket / JSON gateway: same method names and parameters as the Python SDK, with a bundled JS client and demo pageYour own tools and UIs, remote bots

Start with one sentence

Setup takes about 15 minutes. Leave the rest to your agent.