Your model does the thinking, War3AI gets it done
For an agent to get things done on its own in a real environment, it needs three things: actions it can state clearly, feedback it can understand, and an environment where mistakes are safe. We turned Warcraft III into that environment.
Why agents can get things done here
The usual way to make an AI play a game is to simulate the mouse and keyboard and recognize what's on screen, and then guess whether anything worked. We went the other way and designed every API around one question — what does a model need in order to correct itself?
Say only what to do
Move, attack, build, cast… Semantic commands map directly to player actions; units use four-character codes and abilities use order strings, the same names the game uses. The model needs no low-level knowledge.
Structured feedback for every action
The receipt comes back in the same frame, with whether it was accepted, the reason code, the order before and after, and the execution time. “Why it didn't work” is a number, not a guess.
No need to look at the screen
Snapshots show how things stand now; the event stream shows what happened: every hit, every kill, every completed production. Agents can check their own judgment.
The docs are the API
api.json is generated from code, and each of the 103 methods is labeled with its test status and mechanism; llms.txt covers the whole site in one read. The model has no reason to invent methods.
Failure is safe
If the bot throws an exception, the next tick carries on; commands past their deadline never execute; if the bot process crashes, the game keeps running. Agents can experiment freely.
Fair from day one
Fair mode exposes only what's within vision, matching the Arena rules. What an agent learns locally carries over to the Arena without changing a line.
Agent writes the bot
The fastest way to get started, no programming required. Give the manual, api.json and an example to Claude Code, Cursor, Codex or any chat model, and describe your strategy in plain language. It writes a Python file that subclasses openwar3.Bot, and one command puts it in a game.
- The prompt template includes 15 “rules learned in live games”
- It must use only methods in api.json, so nothing gets invented
- Paste errors back verbatim, describe what you see, and have it fix them
Sure. I'll follow the structure of rush_bot.py and use only methods from api.json:
- Only
idle_workers()get sent to gather, 5 per mine - Keep 1 unit in the barracks queue; watch
production().blockedfor food blocks - Pick the expansion with
path_distanceandattack_movethere movethe hero home below 30% HP;reviveit if it dies
Let the agent iterate on its own
A coding agent that can run commands can also take over playing a game and checking the results. Give it an unattended command, a structured match report and a clear goal, and it will improve the bot round after round.
- Run
play.py --speed 200 --minutes 10 --fair - Read the match report: production timeline, losses, top rejection reasons
- Change the one thing that matters most, noting the reasoning in a comment
- Run again, and stop to report after several games with no improvement
{
"reason": "到时间了",
"clock": 612.4,
"timeline": [
[15, "training", "hpea"],
[118, "construction", "hbar"],
[161, "training", "Hamg"],
[402, "upgrade", "hkee"]
],
"lost": {"hfoo": 7, "Hamg": 1},
"killed": {"ogru": 4, "ngno": 6},
"rejects": [
["train hfoo: rejected(人口不够)", 41],
["attack ogru: rejected(目标看不见…)", 9]
]
}The top rejection is “not enough food,” 41 times — the barracks kept waiting on food. Per rule 12, build a farm as soon as production().blocked shows up instead of waiting until food is nearly capped. This round changes only that one thing.
LLM as strategy coach
“What to save for, where to put workers, whether to fight or hold this minute” — big-picture calls like these, which weigh priorities against each other, clash when written as if / else, yet they're exactly what LLMs are good at. The reference brain asks a local model every 20 game seconds and gets back strict JSON.
{
"diagnosis": "Gold 820, lumber 60: lumber is the bottleneck and the barracks stopped for lack of it",
"workers": {"gold": 10, "lumber": 7},
"priority": ["hpea", "hhou", "hfoo"],
"posture": "hold",
"avoid": ["Don't research Iron Plating first when lumber is short"]
}Worker split 10 / 7 · training priority hpea → hhou → hfoo · stance hold
Issues orders, reads receipts, micros
Let units talk
Speech bubbles pop up over any unit, as any character. Multiple units can talk at once, and every bubble's style can be customized. Hook up a local LLM, and one line goes in while the reply streams out over the unit's head. Peasant break-room banter, hero dialogue and battle commentary are all ready to go.
POST http://127.0.0.1:8872/api/chat
{
"inst": 16, "unit": "0x14A12614", "name": "Mountain King",
"persona": "You are the Mountain King: boisterous, fond of ale, one or two casual sentences",
"message": "There's a pack of ogres up ahead. Do we charge?",
"stream": true
}
→ {"reply": "Charge! Just let me finish this ale!",
"first_token_ms": 283, "total_ms": 342}AI companion
Bring your own AI partner into RPG and custom maps. It follows you, helps you fight, heals you when you're low and chats with you when things are quiet — its lines can come from a local LLM. Subclass one class, change a few attributes, and it's your own companion.
Checked in order every tick; the first rule that applies wins
- Retreat When it's low on HP with enemies nearby, it falls back behind you
- Heal When your HP is low and the spell is ready, it heals you
- Assist It hits whatever is attacking you first, then whatever you're attacking
- Follow It catches up when too far away, and runs straight back past a certain distance
- Chat When there's no fighting, it says something every minute or two
Type -follow / -stay / -heal / -hi in chat, or right-click it
LLM calls tools directly
MCP-capable clients — Claude Code, Claude Desktop, agent frameworks for local models — hook up war3_mcp.py, and the LLM can read the game, issue commands, talk to the player on screen, ask the player with pop-up cards and look at screenshots. No code to write first; whatever you think of, just have it do it.
- 10 tools: one-page overview, units, events, call any public API, look up APIs, screen toast, overhead speech, ask the player, screenshot, JASS
- Three roles: dev, player (commands one player only and sees only its vision), observer (read-only)
- It connects to the game only on the first tool call, so the game can be started later
Called 2 tools:
war3_overview→ gold 500 · food 10/12 · town hall 1 · peasants 5 · Paladin 1war3_ask_player→ three cards in the middle of the screen; the game pauses while you choose- You clicked “Expand”; the result came straight back into the conversation, and I'm assigning peasants next
The agent takes the field
The Arena speaks WebSocket / JSON: each tick the referee sends an observation filtered by vision, and the bot replies with a set of actions. Any language and any model can connect — even with no code at all, by having an LLM emit JSON tick by tick. The single-machine version works today: the gateway's player role lets you command just one player and see only its vision; what the Arena still needs is a referee everyone trusts.
- Ticks run on game time; a slow side only hurts itself, never the other
- Every action is checked for unit ownership first and recorded for replay
- Bots written with the SDK can play by swapping
GameforArenaGame
{"t": "obs", "tick": 57, "gameMs": 11400, "me": 1,
"res": {"gold": 320, "lumber": 150, "food": [18, 30]},
"units": [
{"id": 101, "type": "hfoo", "owner": 1,
"x": -4500, "y": 2200, "hp": 380, "hpMax": 420}],
"visibleEnemies": [
{"id": 733, "type": "ogru", "owner": 2,
"x": -3900, "y": 2500, "hp": 700}],
"deadlineMs": 180} id is a stable ID assigned by the referee that stays the same for the whole game. {"t": "act", "tick": 57, "actions": [
{"do": "attack", "unit": 101, "target": 733},
{"do": "train", "unit": 5, "code": "hfoo"},
{"do": "cast", "unit": 7, "spell": "thunderbolt", "target": 733},
{"do": "move", "unit": 102, "x": -5000, "y": 2000}]} category == "command", with the same parameter names; a late act counts as a pass. Hand these to your agent
All plain text or JSON — no login, no rendering. Agents can fetch them directly.
| What you use | How to connect | Best for |
|---|---|---|
| Coding agents such as Claude Code / Cursor / Codex | Open it in the repo directory and have it read the manual, api.json and the examples; it can run games, read reports and iterate on its own | Writing bots, autonomous iteration |
| Chat models with web access | Have it read war3ai.com/en/llms-full.txt first | Writing bots |
| Chat models without web access | Paste the manual, api.json and an example into the prompt | Writing bots |
| Local models (LM Studio / Ollama) | OpenAI-compatible API; MoE models with thinking turned off are recommended | Coaching, voice lines (latency-sensitive) |
| Cloud APIs (Claude / GPT / Gemini / DeepSeek…) | OpenAI-compatible or each vendor's SDK; the coaching layer is async, so a few seconds of latency is fine | Coaching; direct play in the future |
| MCP-capable clients (Claude Code / Claude Desktop…) | Hook up tools/war3_mcp.py: 10 tools to read the game, issue commands, ask the player and take screenshots directly | Commanding as you play, play buddy, commentary |
| Any language / browser / another machine | WebSocket / JSON gateway: same method names and parameters as the Python SDK, with a bundled JS client and demo page | Your own tools and UIs, remote bots |
Start with one sentence
Setup takes about 15 minutes. Leave the rest to your agent.