Agent Self-Iteration
Let a coding agent play games, read the results, change the code and play again on its own. It needs a command that runs unattended, a structured game report, and a clear goal.
In Write a Bot with an LLM, the “play a game → observe → tell the model” step is done by you. A coding agent that can run commands (Claude Code, Codex, Cursor’s agent mode, etc.) can take over that step too, closing the loop:
change code ──► play a game (unattended) ──► read game report ──► find the one thing that matters most ──┐
▲ │
└──────────────────────────────────────────────────────────────────────────────────────────────────────┘
For this loop to actually converge, the agent needs three things.
1. A command that runs unattended
python tools/play.py --bot brains/my_bot.py --speed 200 --minutes 10 --fair
--minutesguarantees the game ends (in wall-clock minutes), so the agent never gets stuck in a game;--speed 200uses 2× game speed to save time — but inside the Bot, wait by the game clock (g.clock()), not with a wall-clocksleep;--fairmakes it follow Arena rules from day one: it can only see what’s in vision;- When the run ends, the terminal prints the end reason, e.g.
我方没有单位了(“we have no units left”) or到时间了(“time’s up”); whatever the Bot itselfprints also shows up in the terminal.
The game simulation stops while the window is minimized. Have the agent start the game in the default windowed mode, and make sure it doesn’t use the same instance number as the one you’re using (--inst).
2. A structured game report
Terminal output is for humans. What the agent should read is a JSON file: what happened, what failed, and why. The SDK already gives you all the raw material — receipts carry reason codes, and the event stream carries production completions and casualties. Just collect them:
import collections, json, time
from openwar3 import Bot
class Recorder(Bot):
"""Adds a game report to a Bot. Subclass it, then call super() in your own on_start / on_event."""
def on_start(self, g):
self.rejects = collections.Counter() # "train hfoo: rejected(人口不够)" -> count (= not enough food)
self.timeline = [] # [game seconds, category, four-character code]: train / research / build / upgrade completed
self.lost = collections.Counter() # what we lost
self.killed = collections.Counter() # what we killed
def check(self, r, what):
"""Wrap a command to record rejection reasons: self.check(g.train(b, "hfoo"), "train hfoo")"""
if r is not None and not r:
self.rejects[f"{what}: {r.reason}"] += 1
return r
def on_event(self, g, ev):
me = g.me()
if ev.kind == "production.done" and ev.owner == me:
self.timeline.append([round(ev.clock), ev.done_kind, ev.done_code])
elif ev.kind == "unit.died":
(self.lost if ev.owner == me else self.killed)[ev.type] += 1
def on_end(self, g, reason):
report = {"reason": reason, "timeline": self.timeline, "lost": self.lost,
"killed": self.killed, "rejects": self.rejects.most_common(10)}
try: # the game may already have exited; skip it if it can't be read
report |= {"clock": g.clock(), "resources": g.resources(),
"army": len(g.my_army()), "workers": len(g.my_workers())}
except Exception:
pass
with open(f"run_{int(time.time())}.json", "w", encoding="utf-8") as f:
json.dump(report, f, ensure_ascii=False, indent=1)
Questions this report can answer:
| Signal | Where it comes from | What it tells you |
|---|---|---|
| Most frequent rejection reasons | Receipt reason / verdict | Constantly food blocked (3), ordering things you can’t afford (8 / 9), attacking targets in the fog of war (1001), training a hero that already died (221) |
| Production timeline | production.done events (with game seconds taken) | When the first hero came out, when you tiered up, whether the Barracks kept producing; compare with pro players’ openings |
| Casualties on both sides | unit.died events | Whether you keep feeding units, how many times the hero died, whether creeping paid off |
| End reason | on_end(g, reason) | 我方没有单位了 (“we have no units left”) = loss; 到时间了 (“time’s up”) = no winner yet |
| Final army and resources | One snapshot read at on_end | Gold piling up unspent = production can’t keep up; too few workers = the economy never took off |
Programmatic win/loss detection is one of the Arena’s foundation experiments and is still on the roadmap. For now, you can treat 我方没有单位了 (“we have no units left”) as a loss, and approximate a win as “every visible enemy building is destroyed”.
3. A clear goal and a few constraints
Give the agent the following, adjusted to your goal:
Goal: make brains/my_bot.py reliably beat the Easy computer on Echo Isles (Human vs random race).
Each round:
1. Run python tools/play.py --bot brains/my_bot.py --speed 200 --minutes 10 --fair
2. Read the terminal output and the latest run_*.json: end reason, production timeline, most frequent rejection reasons, casualties on both sides
3. Find the ONE problem that affects the result most, and change only that; write the reason for the change and the data behind it in a code comment
4. Go back to step 1. If there's no improvement for 3 games in a row, stop and tell me the report and your assessment
Constraints:
- Only use methods in docs/api.json; don't invent APIs
- Don't re-issue the same command to the same unit every tick; only order idle units
- Keep --fair (only use enemies visible in vision)
- Before changing code, run python tools/run_tests.py to make sure the examples aren't broken
Habits that help the loop converge faster
- Change one thing at a time. If you change three things at once and win, you don’t know which one helped; if you lose, you don’t know which one broke it.
- Play enough games to compare. The same matchup has a lot of randomness; two games only reveal very large differences. Judge “did it improve” by the trend over at least several games.
- Fix rejections before tuning strategy. The most frequent rejection reason in the receipts is often the Bot’s biggest bug.
- Write your reasoning into comments. The next round’s agent (or the next conversation) can read from the comments why the code is the way it is, and won’t revert fixes.
- Offline tests as a safety net. Write unit tests for key logic that don’t need a running game (the example Bots’ tests are in
brains/examples/tests/), and have the agent run them after every change.