# Speech bubbles and local models

> Pop up a speech bubble over any unit in the game, speaking as any character. Hook up a local LLM, and one line goes in while the reply appears over the unit's head.

Source: https://war3ai.com/en/docs/speech/

Bubbles are a presentation layer: they don't affect who wins, and they're great for streaming, commentary and debugging.

- Any unit can speak as any character, and multiple units can talk at once;
- Each bubble's font size, color, width, tail, opacity and typing speed can be customized individually;
- Connects directly to a local LLM (LM Studio) with streaming output: the bubble updates as the text is generated.

## Using it from a bot

The simplest way is the SDK's built-in `say`:

```python
g.say(hero, "With me! Charge!", seconds=4)
```

## Starting it and the UI

**Easiest: Farsight's home page, Control Center.** First click **Local LLM → Start and load model** (starts the LM Studio local server and loads the configured model into VRAM), then **Chat bubbles → Start**. The cards also let you view logs, stop and restart.

The UI is the **Chat Bubbles** page in Farsight's left sidebar: make units speak (pick a unit, write the text, adjust the style, chat with the model), peasant break room, camera dialogue, battle triggers and model settings, all acting on the instance selected in the top bar.

You can also use the command line:

```bash
python speech/speak_launch.py              # start the local model server + load and warm up the model + start the bubble API
python speech/speak_launch.py --restart    # restart the API after changing code
python speech/speak_launch.py --stop       # stop the API and unload the model from VRAM
```

Every step is skipped if it's already running, so running it again has no side effects.

## HTTP API

The default is `http://127.0.0.1:8872/` (the port is `ports.speech` in `openwar3.json`), and any program can call it.

### Make a unit speak: `POST /api/say`

```json
{
  "inst": 16,
  "bubbles": [
    { "unit": "0x14A12614", "name": "Mountain King", "text": "With me! Charge!" },
    { "unit": "0x14A12924", "name": "Archmage", "text": "I'll cast Blizzard.",
      "style": { "font_px": 26, "text_color": "#FFE080", "bg_color": "#C0102040" } },
    { "screen": [960, 110], "key": 1, "name": "Narrator", "text": "The first orc wave arrives in 30 seconds.",
      "style": { "tail": false, "type_ms": 0 } },
    { "world": [-4684, 2644], "key": 2, "text": "Rally point", "style": { "font_px": 16 } }
  ]
}
```

| Field | Description |
|---|---|
| `unit` / `world` / `screen` | Pick one: follow a unit (sits right above its health bar when it has one) / map coordinates / screen pixels (for narration) |
| `name` | Speaker shown on the first line; anything you like, it doesn't have to be the unit |
| `text` | Body text, wrapped automatically |
| `duration_ms` | How long to show it; 0 = automatic, 3–5 seconds |
| `key` | ID for world / screen bubbles; a new message with the same key replaces the old one |
| `update` | If the same bubble already exists, only swap the text without resetting the timer (for streaming) |
| `style` | `font_px`, `max_width_px`, `text_color`, `bg_color`, `border_color`, `tail`, `side_px`, `opacity`, `type_ms`, `font`… |

Up to 32 bubbles at once; the per-frame cost averages about 0.1–0.2 ms.

### Chat with a local model: `POST /api/chat`

```json
{
  "inst": 16, "unit": "0x14A12614", "name": "Mountain King",
  "persona": "You are Muradin, the Mountain King from Warcraft: boisterous and fond of ale. Reply in one or two casual sentences, under 40 words.",
  "message": "There's a pack of ogres up ahead. Do we charge?",
  "stream": true
}
```

It returns `{"reply": "...", "first_token_ms": 283, "total_ms": 342}`, and by then the reply is already showing over that unit's head. Each unit remembers its last 6 rounds of conversation.

### Other endpoints

| Endpoint | Description |
|---|---|
| `GET /api/instances` | Running games |
| `GET /api/units?inst=16&mine=true&heroes=true` | Unit list (with Chinese display names, coordinates, HP) |
| `POST /api/clear` | Clear one bubble or all of them |
| `GET /api/llm`, `POST /api/llm` | View / change the model config (`base_url`, `model`, `max_tokens`, `temperature`) |
| `POST /api/banter` | Peasant break room: workers at home take turns griping in character, with an opening roll call (every battle stat is real data) |
| `POST /api/camtalk` | Camera dialogue: heroes and followers on camera talk to each other in character |
| `POST /api/events` | Battle triggers: fight starts, fight ends, hero dies, tier up, building destroyed… it only speaks when something happens |

## Choosing a local model

Measured on a single RTX 5090 (5 in-game lines):

| Model | VRAM | Speed | One reply | Verdict |
|---|---|---|---|---|
| **Qwen3.6-35B-A3B** (MoE, only 3B active at a time), Q4, thinking off | 20.6 GB | ~142 tokens/s | **~0.3 s** (first token ~0.27 s) | Recommended: fast, with natural Chinese role-play |
| gpt-oss-20b (MXFP4), reasoning low | 11.3 GB | ~280 tokens/s | 0.3–0.8 s | Use when VRAM is tight; Chinese is a bit flat |
| Qwen3.6-27B (dense), Q4 | 17.2 GB | ~39 tokens/s | Still thinking at 5.5 s | Not suitable for real-time dialogue |

- **Speed depends on the parameters active per step, not the total**: the 35B MoE activates only 3B and is 3–4× faster than the 27B dense model.
- **You must turn off "thinking"**: otherwise every token goes into thinking and not a single word of reply comes back.
- Bubbles type out at about 22 characters per second, so generation speed is no longer the bottleneck; what really shapes the experience is **time to first token**.

> **Make the lines ring true**
>
> Always feed the model real battle data (games played, wins and losses, army size, resources in the bank), and state explicitly that it may use only these facts. In testing, without that constraint the model made up battles that never happened.
