# LLM as Strategy Coach

> Hand "what to save for, where to put workers, attack or hold this minute" to an LLM, and let the rule layer only execute and veto. The reference brain already works this way; this page explains the pattern and the pitfalls.

Source: https://war3ai.com/en/docs/llm-coach/

Once your Bot grows past a certain point, you'll notice the economy rules are stacked one on top of another: a rule for how many lumberjacks, a rule for 5 workers per mine, a rule to halve lumber when there's too much, a rule to send more to gold when gold is short and lumber is plentiful… Each rule is correct on its own, yet together they produce situations like "the mine is short of workers while every peasant is chopping trees" — situations **no single rule is responsible for**.

Judgments like "look at the whole picture and set priorities" were never a good fit for `if / else`, but they're exactly what LLMs are good at. The reference brain (`brains/xwar3/strategy/brain/coach.py`) uses the layering below.

## Layers

```text
LLM (advisor)           Once every 20 game seconds, async, never blocks a tick
  Input: a one-page snapshot of the game (resources, food, worker distribution, mines, unit types, tech, heroes, enemy intel, recent events)
  Output: strict JSON — a one-line diagnosis + worker split + what to build first + this minute's posture + things to avoid
        │
        ▼  whitelist + min/max clamping + veto
Rule layer (Bot, every tick)  Translates advice into "biases" on existing abilities: worker split, build / train priority, attack posture
        │
        ▼
Execution layer (SDK / reflex layer)  Issues orders, reads receipts, micro
```

## Output contract

Make the model output JSON with a fixed set of fields — no additions, no omissions:

```json
{
  "diagnosis": "One sentence: the biggest problem in the game, which must be backed by the input data",
  "workers":   { "gold": 10, "lumber": 6 },
  "priority":  ["hpea", "hhou", "hbar"],
  "posture":   "creep",
  "avoid":     ["Don't research Iron Plating first when lumber is short"]
}
```

| Field | How the rule layer uses it | Reference brain's clamp |
|---|---|---|
| `workers` | Target number of workers on gold and on lumber | Gold 2 ~ 25, lumber 1 ~ 20; the sum can't exceed the total number of peasants |
| `priority` | Priority order for training / building / research | At most 4; only four-character codes that appear in the "allowed codes" table are accepted |
| `posture` | This minute's posture | Must be one of `attack` `defend` `creep` `expand` `recover` `hold` |
| `avoid` | Things not to do this minute | At most 2 |
| `diagnosis` | Only used for logs and the console display | — |

Use one prompt per race, covering only that race's specific trade-offs (Human's cooperative building and Militia, Orc's Burrows, Undead's Haunted Gold Mine, Night Elf's Entangled Gold Mine…). Put the common rules in a shared section — don't copy them four times.

## Four hard constraints

The reference brain learned each of these the hard way:

1. **The advisor never issues unit orders directly.** It can't see what's happening on a 150 ms timescale, and it hallucinates. It only changes targets and priorities; who goes where and who attacks what is still decided by the rule layer and the reflex layer — command authority can only have one owner.
2. **Async.** One advisor call takes about 1 second and runs on a background thread; the latest result wins, and it **never blocks a tick**. If the model isn't up, times out, or answers garbage, act as if this layer doesn't exist and fall back to pure rules. Advice that's too old (older than 3 intervals) isn't used either.
3. **Whitelist + clamping.** Every field must map to an existing ability, and numeric values are clamped to a sensible range. Anything unrecognized is **counted and then discarded**, not silently ignored.
4. **Count everything.** How many times you asked, how many succeeded, how many timed out, how many were clamped, how many times each field was adopted — publish all of it along with the last input sent to the model. Otherwise "is this layer actually doing anything?" is a question you can't answer.

> **Things that degrade safely are the easiest to degrade silently**
>
> The advisor is designed so that "failure = act as if this layer doesn't exist", so when the model service isn't running, the Bot behaves exactly like pure rules and nothing looks wrong from the outside. The reference brain once went a whole day with the advisor unable to connect on all 6 instances, and nobody noticed. Always publish "last successful call time" and "last failure reason" — that's exactly what the "strategy coach" page in the [Farsight console](https://war3ai.com/en/docs/console/) is for.

## Implementing it in your own Bot

Below is a minimal skeleton that works with any OpenAI-compatible API (LM Studio, Ollama, or a cloud API) and uses only the standard library:

```python title="coached_bot.py"
import collections, json, threading, urllib.request
from openwar3 import Bot

BASE = "http://127.0.0.1:1234/v1"            # LM Studio / Ollama / any OpenAI-compatible service
MODEL = "your-model"
POSTURES = {"attack", "defend", "creep", "expand", "recover", "hold"}
SYSTEM = """You are a Warcraft III macro coach. You handle only economy and strategy, not micro.
Output JSON only, with fixed fields: {"diagnosis": one sentence, "workers": {"gold": integer, "lumber": integer},
"priority": [four-character codes, at most 4, only from allowed], "posture": one of six, "avoid": [at most 2]}
Base everything on the game data you are given; don't make up anything that isn't in the data."""

def ask(state: dict) -> dict:
    body = {"model": MODEL, "temperature": 0.3, "max_tokens": 260,
            "messages": [{"role": "system", "content": SYSTEM},
                         {"role": "user", "content": json.dumps(state, ensure_ascii=False)}]}
    req = urllib.request.Request(f"{BASE}/chat/completions", json.dumps(body).encode(),
                                 {"Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=8) as r:
        text = json.load(r)["choices"][0]["message"]["content"]
    return json.loads(text[text.index("{"): text.rindex("}") + 1])

class CoachedBot(Bot):
    EVERY = 20.0                              # game seconds: macro decisions happen on a scale of minutes, no need to ask every tick
    allowed = {"hpea", "hfoo", "hrif", "hkni", "hhou", "hbar", "hbla", "Rhme", "Rhar"}

    def on_start(self, g):
        self.plan, self.plan_at, self.asked_at, self.busy = {}, -1e9, -1e9, False
        self.stats = collections.Counter()

    def summary(self, g) -> dict:             # read the snapshot on the main thread; the background thread never touches g
        res = g.resources() or {}
        return {"clock": round(g.clock() or 0), "gold": res.get("gold"), "lumber": res.get("lumber"),
                "food": [res.get("food_used"), res.get("food_cap")],
                "workers": len(g.my_workers()), "idle_workers": len(g.idle_workers()),
                "army": collections.Counter(u.type for u in g.my_army()),
                "enemy_seen": collections.Counter(u.type for u, _t, _age in g.last_seen(max_age=90)),
                "night": g.is_night(), "allowed": sorted(self.allowed)}

    def consult(self, state, now):
        try:
            p = ask(state)
            self.stats["ok"] += 1
            posture = p.get("posture")
            if posture not in POSTURES:
                self.stats["bad_posture"] += 1                       # count, then discard — never silently
                posture = "hold"
            self.plan = {"gold": min(25, max(2, int(p["workers"]["gold"]))),      # clamp
                         "lumber": min(20, max(1, int(p["workers"]["lumber"]))),
                         "priority": [c for c in p.get("priority", []) if c in self.allowed][:4],
                         "posture": posture}
            self.plan_at = now
        except Exception as e:                                       # timeout / garbage answer: act as if this layer doesn't exist
            self.stats[f"error:{type(e).__name__}"] += 1
        finally:
            self.busy = False

    def on_tick(self, g):
        now = g.clock() or 0.0
        if not self.busy and now - self.asked_at >= self.EVERY:
            self.busy, self.asked_at = True, now
            threading.Thread(target=self.consult, args=(self.summary(g), now), daemon=True).start()
        plan = self.plan if now - self.plan_at <= 3 * self.EVERY else {}   # don't use advice that's too old
        # ↓ rule layer: with an empty plan, follow the default rules; with a plan, only adjust the split, priorities and posture — the actual orders are still decided by rules
        ...
```

## Choosing a model

| Situation | Recommendation |
|---|---|
| Local, needs to be fast | MoE models (which activate only a small fraction of parameters per call) are much faster than dense models of the same size. The reference brain uses Qwen3.6-35B-A3B (LM Studio, Q4): median **1.09 s**, slowest 1.45 s, and 5/5 outputs parse directly with `json.loads` |
| Local "thinking" models | **You must turn off the thinking section**, otherwise all the tokens go to thinking and no JSON ever comes out. LM Studio ignores `/no_think`; the reference brain switched to `/v1/completions`, builds the ChatML itself, and pre-fills an empty `<think></think>` followed by a `{` |
| Cloud models | Latency is usually higher, but this layering is async by design; macro decisions are measured in minutes, so a few seconds of latency is fine |

> **Note**
>
> The same model can also voice your units: see [Speech bubbles and local models](https://war3ai.com/en/docs/speech/). If you want the model to issue orders directly every tick (instead of acting as an advisor), wait for the [Arena](https://war3ai.com/en/arena/) JSON gateway.
