Join the three evidence sources that were kept apart, and extract the useful
result: each source independently decides what belongs in code and what belongs
to the model, and every serious bug we have found broke one of those rules.
* Game mechanism -> arithmetic and rules belong in CODE
* Internal state -> anything directly readable or writable belongs in CODE
* TypeSafe -> judgements go to the MODEL, in the primitive matching the
answer's shape
* never blocked in boss fights -> game mechanism (arithmetic left to the model)
* card quality asked as a Noul -> primitives (a spectrum forced into yes/no)
* indices read from the wrong array -> internal state
The counter-example is kept too: the Score we tried for fight-level planning was
unusable, so the rule is "match the primitive to the shape, and verify" rather
than "use Score more".
Also records the ordered plan (measurement first, then re-run, then primitives,
then deck composition), the continual-learning design, and why it is blocked:
with 0 wins in 37 runs the reward has no gradient, so training would fit the bug
rather than the game.
335 lines
15 KiB
Markdown
335 lines
15 KiB
Markdown
# 10 — Synthesis: game, state and primitives → next iteration
|
||
|
||
Status: **forward-looking**. This doc is the join of three evidence sources that
|
||
until now were kept apart:
|
||
|
||
| Source | Doc | What it gives us |
|
||
|---|---|---|
|
||
| STS2 game mechanism | `01`, `08` | What actually decides a run |
|
||
| Internal game state we control | `04`, `08` | What is observable, and what we can change |
|
||
| TypeSafe / System One primitives | `02`, `09` | What shape to ask the model |
|
||
|
||
It records the design conclusions and the planned iteration. Nothing here is a
|
||
measurement unless it says so; the measurements live in `05`, `07`, `08`, `09`.
|
||
|
||
---
|
||
|
||
## 1. The synthesis in one line
|
||
|
||
**Each source independently decides what belongs in code and what belongs to the
|
||
model — and every serious bug we have found is a violation of one of those three
|
||
rules.**
|
||
|
||
| Source | Rule it imposes |
|
||
|---|---|
|
||
| Game mechanism | Anything that is arithmetic or a rule belongs in **code**. HP, block, lethal, damage projection, energy. |
|
||
| Internal state | Anything we can read or write directly belongs in **code**. Index spaces, save files, prefs, action enumerations. |
|
||
| TypeSafe primitives | Anything that is a *judgement over unstructured meaning* goes to the **model** — and must use the primitive that matches the answer's shape. |
|
||
|
||
### The three bugs, mapped to the three rules
|
||
|
||
| Bug | Rule broken | Doc |
|
||
|---|---|---|
|
||
| Bot never blocked in boss fights; 9 runs lost to `THE_KIN_BOSS` | **Game mechanism.** Blocking is arithmetic (`projected_incoming > affordable_loss`) and we left it to a damage-biased model question. | `08` |
|
||
| Card quality asked as a `Noul` and `argmax`'d | **Primitives.** A spectrum judgment forced into yes/no. `primitives.md`: "A Noul value of 0.5 means the model gives yes and no equal probability. It does not mean the candidate has a medium skill level." | `09`, §2 |
|
||
| Card indices read from the wrong array (5 spaces, similar names) | **Internal state.** We misread what the index pointed at. | `04` |
|
||
|
||
That mapping is the useful part. It turns three unrelated bug hunts into one
|
||
checkable rule, and it predicts where the next bug is: **any place we ask the
|
||
model something that is really arithmetic, or use the wrong primitive for the
|
||
answer shape.**
|
||
|
||
### The counter-example that keeps us honest
|
||
|
||
`09` §3 records a `Score` we tried for fight-level planning ("Race / Trade /
|
||
Control") that was **unusable on 4 of 6 states**, confidence as low as 0.01.
|
||
So the rule is not "use Score more". It is "match the primitive to the answer
|
||
shape, and verify the shape works before adopting it."
|
||
|
||
---
|
||
|
||
## 2. What the game actually rewards (from `08`)
|
||
|
||
Established by mining 37 run files, not by reading guides:
|
||
|
||
- **81% of runs (30/37) die in Act 1.** Deck composition is downstream of that.
|
||
- **43% die to an Act 1 boss; `THE_KIN_BOSS` alone kills 9.**
|
||
- The bot is **not** bloating, descriptively: 23% of rewards were skipped (249 of
|
||
323 rewards took a card), one card per reward at most. The per-card rate
|
||
(249/1087 = 23%) is mechanically diluted by offer size and does not diagnose
|
||
skip policy; we have no STS2 benchmark for a good skip rate.
|
||
- The Kin fights lasted **5–10 turns** and cost **44–80 HP**, ~10–13 a turn,
|
||
with block cards in hand. That is a defense failure, and it is now fixed.
|
||
|
||
**Consequence for planning:** any work on archetypes, synergy scoring or deck
|
||
"lean weight" is tuning a system the bot does not survive long enough to use.
|
||
The order is: survive Act 1 → then optimise the deck.
|
||
|
||
---
|
||
|
||
## 3. What the state gives us that we are not yet using
|
||
|
||
Three levers found but not yet exploited:
|
||
|
||
### 3.1 The run files are a labelled dataset
|
||
|
||
`history/*.run` records, per map point: every card offered with `was_picked`,
|
||
`damage_taken`, `potion_used`, `rest_site_choices`, `upgraded_cards`, and the
|
||
final deck. That is a supervised dataset of our own drafting and combat, and it
|
||
already invalidated one hypothesis (bloat) and found one bug (defense).
|
||
|
||
**Not yet used for:** labelling decisions, or measuring the gates.
|
||
|
||
### 3.2 Instant Mode is a one-line lever
|
||
|
||
`prefs.save` holds `"fast_mode"` as plain JSON, and the mod's Instant Mode
|
||
checkbox is only `PrefsSave.FastMode = FastModeType.Instant`. Setting it to
|
||
`"instant"` removes animations and should roughly halve wall-clock per run.
|
||
**Needs a game restart** — prefs load at startup.
|
||
|
||
This matters because the binding cost of every experiment is wall-clock. At
|
||
~10 minutes per run, an A/B of 6v6 is two hours.
|
||
|
||
### 3.3 Model answers are now logged with the decision (not in a side trace)
|
||
|
||
**Corrected 2026-09-22.** This section previously said
|
||
`JEV_TRACE=capture/jev_trace.jsonl` was "the entire change". It is not, and it
|
||
was never sufficient on its own: `jev.py`'s trace records a `HH:MM:SS` stamp,
|
||
the questions and the answers — **no run, session, step, state or action**. Two
|
||
calls in the same second are indistinguishable, so the file cannot be joined to
|
||
a decision or to an outcome, which is the only thing a label needs.
|
||
|
||
What the decision log carries now, per decided action
|
||
(`run.py:decision_record`):
|
||
|
||
```json
|
||
{"session":"20260922T034419","step":55,"state_type":"monster",
|
||
"run":{"act":1,"floor":5},"action":"play_card","source":"jev",
|
||
"params":{"card_index":1,"target":"SHRINKER_BEETLE_0"},
|
||
"confidence":0.85,"reason":"jev chose Dismantle","error":null,
|
||
"jev":{"model":"jev-1.13.0","latency_s":0.71,
|
||
"questions":{...},"answers":{
|
||
"best_play":{"kind":"choice","choice":"card1","margin":0.79,"gated":true,
|
||
"probabilities":{"card0":0.01,"card1":0.89,...}},
|
||
"use_potion":{"kind":"noul","noul":0.41,"yes":false},
|
||
"which_potion":{"kind":"choice","choice":"potion0","margin":1.0,"gated":true}}}}
|
||
```
|
||
|
||
- `session` is a unique id per process (`<timestamp>-<6 hex>`). It is the join
|
||
key only because three places now carry it:
|
||
1. every decision row in `capture/decisions.jsonl`;
|
||
2. `capture/sessions.jsonl`, written by `run.py` at exit (via `atexit`, so it
|
||
covers every exit path) and mapping the session to the history run file(s)
|
||
it produced, with `win`, `killed_by`, `seed`, `deck_size`, `map_points`;
|
||
3. the A/B attribution row — `policy \t run \t code-hash \t session` — and the
|
||
report prints each arm's session ids.
|
||
The game's own run file does **not** know our session id and the harness
|
||
originally mapped only a policy to a filename, so without (2) and (3) nothing
|
||
lined up and there were still no labels.
|
||
- `jev` is `null` when the decision asked nothing (code paths, fallbacks), which
|
||
is a fact about the decision rather than a gap.
|
||
- `RecordingClient` in `run.py` captures the answers, so nothing in `brain.py`
|
||
had to change.
|
||
|
||
`JEV_TRACE` still works and is still worth setting, but it is now redundant for
|
||
labelling. It is no longer the prerequisite.
|
||
|
||
---
|
||
|
||
## 4. Where our Jev usage stands (from `09`, plus this audit)
|
||
|
||
Measured usage of the documented surface:
|
||
|
||
| Capability | Used |
|
||
|---|---|
|
||
| `Choice` | **7** questions — `best_play`, `which_potion`, `target`, `next_node`, `best_option`, `best_relic`, `best_bundle` |
|
||
| `Noul` | **6** question families — `use_potion`, `want_any`, `skip_all`, `good_card{i}`, `good_relic{i}`, `good_card{i}` (card_select), `worth_item{i}` |
|
||
| `Score` | **0** — though `score()` is defined at `jev.py:169` |
|
||
| Structured criteria (`entry()`) | **0** |
|
||
| Contrastive Noul criteria (`true=`/`false=`) | **0** |
|
||
| `focus=` / `inspect=` / `compare=` | **0** |
|
||
| `gate_choice()` / `margin()` | **0** in `brain.py` (`gate()` is used; `margin()` only inside `jev.py`) |
|
||
|
||
Every question is a bare one-line string. The docs permit that only when the
|
||
question is *short and unambiguous*; ours are neither, e.g.:
|
||
|
||
```python
|
||
noul("Is any relic in `relics` worth taking over skipping?")
|
||
```
|
||
|
||
`09` §2 measured structured criteria on card play: **same card chosen 6/6**,
|
||
margin 0.425 → 0.473. A small effect — worth adopting where disambiguation
|
||
costs us, not as a blanket rewrite. That measurement stands.
|
||
|
||
---
|
||
|
||
## 5. The next iteration, in dependency order
|
||
|
||
Ordered by dependency, not by appeal. Each step states what it unblocks.
|
||
|
||
### Step 1 — Enable measurement (blocks everything else)
|
||
|
||
**Partly landed 2026-09-22.** The answers and the join are in; the corpus is not.
|
||
|
||
Done:
|
||
|
||
- every decision row carries the model's questions and answers, or `null` when
|
||
nothing was asked (`run.py:decision_record`, `RecordingClient`);
|
||
- a unique `session` id on every row, and `capture/sessions.jsonl` at exit
|
||
mapping it to the run file(s) produced, with the outcome fields;
|
||
- `ab_card_skip.sh` stamps the same session into its attribution rows.
|
||
|
||
Still to do:
|
||
|
||
- capture **all** state types, not just combat (`run.py` currently dumps
|
||
`live_{step}_combat.json` only);
|
||
- namespace captures per run so they stop overwriting across sessions;
|
||
- log the *features* the decisions turn on (the `features` block below is a
|
||
sketch, not what is written today — the questions and answers are the
|
||
evidence currently recorded).
|
||
|
||
```json
|
||
{"run_id":"1790...","step":42,"state_type":"card_reward",
|
||
"features":{"card0_power":2.1,"card0_synergy":1.4},
|
||
"action":"select_card_reward:0","source":"jev","confidence":0.71,
|
||
"gate":"pass"}
|
||
```
|
||
|
||
**Why first:** it is the prerequisite for every quality question in `09` and for
|
||
continual learning (§6). With the join in place, the first usable dataset is one
|
||
A/B run away; what is missing now is coverage (all state types, namespaced per
|
||
run) and volume, not mechanism.
|
||
|
||
### Step 2 — Re-run and check whether the ceiling moved
|
||
|
||
The defense fix (`08`, failure #35) is the single highest-leverage change made
|
||
so far. Verify it with the A/B harness, now that `--stop-on-run-end` makes one
|
||
session equal one run:
|
||
|
||
```bash
|
||
./ab_card_skip.sh 6 # 6 runs per arm, one run per session
|
||
```
|
||
|
||
**The question:** does `THE_KIN_BOSS` stop killing 9 of 37? If yes, runs start
|
||
reaching Act 2+ and the reward signal finally varies. If no, the defense fix is
|
||
wrong and nothing downstream is worth building.
|
||
|
||
**Why second:** it is the test of the only substantive finding, and it gates §6.
|
||
|
||
### Step 3 — Match primitives to answer shapes
|
||
|
||
Only after Steps 1–2 give labels:
|
||
|
||
- Card quality: **`Score` with explicit ordered levels**, per card, several
|
||
atomic axes, combined with weights in code (composite scoring). Replaces the
|
||
`Noul` + `argmax` that the docs say is the wrong shape.
|
||
- Structured criteria on the fuzzy Nouls (`worth taking`, `stronger`).
|
||
- Measure each change against the Step 1 trace before keeping it.
|
||
|
||
### Step 4 — Then deck composition
|
||
|
||
Archetypes, synergy, removal priority. Last, because `08` shows the bot does not
|
||
currently reach the point where it matters.
|
||
|
||
---
|
||
|
||
## 6. Continual learning — the design, and the honest blocker
|
||
|
||
The docs give the mechanism directly:
|
||
|
||
> "Combine independent answers with deterministic rules or weighted sums. **For
|
||
> learned composition, use the probabilities as features in a downstream
|
||
> classical machine-learning model.**"
|
||
|
||
So continual learning here is **not fine-tuning Jev**. It is: Jev produces
|
||
features, code combines them, and **the weights are learned from outcomes**.
|
||
|
||
Four stages:
|
||
|
||
1. **Instrument** — Step 1 above.
|
||
2. **Label** — attach the run outcome to each decision row.
|
||
3. **Fit** — a small logistic regression or gradient-boosted tree on
|
||
`features → P(reach Act 2)`. The coefficients *are* the composite weights.
|
||
4. **Close the loop** — retrain periodically and A/B the new weights against the
|
||
old, using the `--stop-on-run-end` harness.
|
||
|
||
The docs also prescribe how to set our gates, which we have so far guessed:
|
||
|
||
> "Test thresholds by plotting confidence against accuracy on your data."
|
||
|
||
`EVENT_TOP_MIN = 0.60`, `EVENT_MARGIN_MIN = 0.30` and `CONFIDENCE_FLOOR = 0.55`
|
||
were chosen by hand. With a trace they can be **measured**. (Note that
|
||
`CONFIDENCE_FLOOR` is inert where it is passed: `gate()` ignores its threshold
|
||
argument for a `ChoiceAnswer`, so the target question is gated by the margin
|
||
rule, not by 0.55.)
|
||
|
||
### The blocker, stated plainly
|
||
|
||
**We have 0 wins in 37 runs.** The top of the reward scale is unobserved, so
|
||
there is nothing to learn from yet. Our only varying signal is map points
|
||
reached — noisy, and confounded by seed luck.
|
||
|
||
Training on this data would fit the bug, not the game. **Step 2 is what creates
|
||
the gradient.** If runs start reaching Act 2 and Act 3, a label appears
|
||
(`P(reach Act 2)`) that actually varies, and Stage 3 becomes possible.
|
||
|
||
---
|
||
|
||
## 7. Refactor implications
|
||
|
||
`brain.py` is ~1700 lines with 13+ handlers, and it now mixes four concerns:
|
||
state parsing, arithmetic, question building, and policy. The three-rule framing
|
||
in §1 suggests the split that the code keeps trying to make:
|
||
|
||
```
|
||
sts2bot/
|
||
game/ card and encounter knowledge, mechanics constants
|
||
state/ observation parsing, the five index spaces, CombatFacts
|
||
policy/ one module per state_type; chooses among enumerated actions
|
||
model/ TypeSafe primitives, a shared question library, the trace
|
||
learn/ decision log, labels, weight fitting
|
||
eval/ A/B harness, metrics, run-file mining
|
||
```
|
||
|
||
Two rules that should hold after the split:
|
||
|
||
1. **`policy/` may not do arithmetic.** If a decision needs a number, the number
|
||
comes from `state/`. This is the rule the defense bug broke.
|
||
2. **`model/` may not be asked a question whose answer is arithmetic.** This is
|
||
the rule the `Noul`-for-spectrum bug broke.
|
||
|
||
Both are testable as structural assertions, in the same style as the existing
|
||
`test_brain.py` hostile-index tests (indices from data, `OFFSET = 5`).
|
||
|
||
**Do not refactor yet.** The measurement in Step 1 should land first: it is
|
||
cheap, it is a prerequisite for everything, and moving files while the corpus
|
||
format is still changing would mean doing it twice.
|
||
|
||
---
|
||
|
||
## 8. Open questions carried forward
|
||
|
||
| Question | Blocked on |
|
||
|---|---|
|
||
| Do the gates (0.45 / 0.60 / 0.30) match their accuracy? | **per-decision labels** — replay, expert judgement, or a code-verifiable ground truth. Run outcomes cannot answer this: one run result attached to a step does not say whether that step's answer was correct. |
|
||
| Is structured criteria worth adopting beyond card play? | trace, then A/B |
|
||
| Does the two-stage cascade beat one broad question? | an optimal-action label |
|
||
| Does map beam search beat the single-step Choice? | `map_point_history` labels |
|
||
| Is `TREMBLE` (offered 45×, taken 0×) correctly rejected, or is that a bug? | a card-quality label |
|
||
| Does the event keyword veto reject benign options? | event captures + outcomes |
|
||
| Is the escalation tier (act / review / escalate) needed? | trace |
|
||
|
||
---
|
||
|
||
## 9. Summary
|
||
|
||
- **Three rules**, one per evidence source, and every serious bug broke one.
|
||
- **One fix landed** (defense) with the highest measured leverage: 9 runs, one
|
||
boss, one cause.
|
||
- **One blocker**: 0 wins means no reward gradient, so continual learning is
|
||
designed but cannot start.
|
||
- **One prerequisite**: enable `JEV_TRACE` and log decisions. Cheap, and it
|
||
unblocks every quality question above.
|
||
- **One test that matters next**: re-run and see whether `THE_KIN_BOSS` deaths
|
||
drop. That single number decides whether the rest of this plan is worth
|
||
building.
|