# 10 — Synthesis: game, state and primitives → next iteration Status: **forward-looking**. This doc is the join of three evidence sources that until now were kept apart: | Source | Doc | What it gives us | |---|---|---| | STS2 game mechanism | `01`, `08` | What actually decides a run | | Internal game state we control | `04`, `08` | What is observable, and what we can change | | TypeSafe / System One primitives | `02`, `09` | What shape to ask the model | It records the design conclusions and the planned iteration. Nothing here is a measurement unless it says so; the measurements live in `05`, `07`, `08`, `09`. --- ## 1. The synthesis in one line **Each source independently decides what belongs in code and what belongs to the model — and every serious bug we have found is a violation of one of those three rules.** | Source | Rule it imposes | |---|---| | Game mechanism | Anything that is arithmetic or a rule belongs in **code**. HP, block, lethal, damage projection, energy. | | Internal state | Anything we can read or write directly belongs in **code**. Index spaces, save files, prefs, action enumerations. | | TypeSafe primitives | Anything that is a *judgement over unstructured meaning* goes to the **model** — and must use the primitive that matches the answer's shape. | ### The three bugs, mapped to the three rules | Bug | Rule broken | Doc | |---|---|---| | Bot never blocked in boss fights; 9 runs lost to `THE_KIN_BOSS` | **Game mechanism.** Blocking is arithmetic (`projected_incoming > affordable_loss`) and we left it to a damage-biased model question. | `08` | | Card quality asked as a `Noul` and `argmax`'d | **Primitives.** A spectrum judgment forced into yes/no. `primitives.md`: "A Noul value of 0.5 means the model gives yes and no equal probability. It does not mean the candidate has a medium skill level." | `09`, §2 | | Card indices read from the wrong array (5 spaces, similar names) | **Internal state.** We misread what the index pointed at. | `04` | That mapping is the useful part. It turns three unrelated bug hunts into one checkable rule, and it predicts where the next bug is: **any place we ask the model something that is really arithmetic, or use the wrong primitive for the answer shape.** ### The counter-example that keeps us honest `09` §3 records a `Score` we tried for fight-level planning ("Race / Trade / Control") that was **unusable on 4 of 6 states**, confidence as low as 0.01. So the rule is not "use Score more". It is "match the primitive to the answer shape, and verify the shape works before adopting it." --- ## 2. What the game actually rewards (from `08`) Established by mining 37 run files, not by reading guides: - **81% of runs (30/37) die in Act 1.** Deck composition is downstream of that. - **43% die to an Act 1 boss; `THE_KIN_BOSS` alone kills 9.** - The bot is **not** bloating, descriptively: 23% of rewards were skipped (249 of 323 rewards took a card), one card per reward at most. The per-card rate (249/1087 = 23%) is mechanically diluted by offer size and does not diagnose skip policy; we have no STS2 benchmark for a good skip rate. - The Kin fights lasted **5–10 turns** and cost **44–80 HP**, ~10–13 a turn, with block cards in hand. That is a defense failure, and it is now fixed. **Consequence for planning:** any work on archetypes, synergy scoring or deck "lean weight" is tuning a system the bot does not survive long enough to use. The order is: survive Act 1 → then optimise the deck. --- ## 3. What the state gives us that we are not yet using Three levers found but not yet exploited: ### 3.1 The run files are a labelled dataset `history/*.run` records, per map point: every card offered with `was_picked`, `damage_taken`, `potion_used`, `rest_site_choices`, `upgraded_cards`, and the final deck. That is a supervised dataset of our own drafting and combat, and it already invalidated one hypothesis (bloat) and found one bug (defense). **Not yet used for:** labelling decisions, or measuring the gates. ### 3.2 Instant Mode is a one-line lever `prefs.save` holds `"fast_mode"` as plain JSON, and the mod's Instant Mode checkbox is only `PrefsSave.FastMode = FastModeType.Instant`. Setting it to `"instant"` removes animations and should roughly halve wall-clock per run. **Needs a game restart** — prefs load at startup. This matters because the binding cost of every experiment is wall-clock. At ~10 minutes per run, an A/B of 6v6 is two hours. ### 3.3 Model answers are now logged with the decision (not in a side trace) **Corrected 2026-09-22.** This section previously said `JEV_TRACE=capture/jev_trace.jsonl` was "the entire change". It is not, and it was never sufficient on its own: `jev.py`'s trace records a `HH:MM:SS` stamp, the questions and the answers — **no run, session, step, state or action**. Two calls in the same second are indistinguishable, so the file cannot be joined to a decision or to an outcome, which is the only thing a label needs. What the decision log carries now, per decided action (`run.py:decision_record`): ```json {"session":"20260922T034419","step":55,"state_type":"monster", "run":{"act":1,"floor":5},"action":"play_card","source":"jev", "params":{"card_index":1,"target":"SHRINKER_BEETLE_0"}, "confidence":0.85,"reason":"jev chose Dismantle","error":null, "jev":{"model":"jev-1.13.0","latency_s":0.71, "questions":{...},"answers":{ "best_play":{"kind":"choice","choice":"card1","margin":0.79,"gated":true, "probabilities":{"card0":0.01,"card1":0.89,...}}, "use_potion":{"kind":"noul","noul":0.41,"yes":false}, "which_potion":{"kind":"choice","choice":"potion0","margin":1.0,"gated":true}}}} ``` - `session` is a unique id per process (`-<6 hex>`). It is the join key only because three places now carry it: 1. every decision row in `capture/decisions.jsonl`; 2. `capture/sessions.jsonl`, written by `run.py` at exit (via `atexit`, so it covers every exit path) and mapping the session to the history run file(s) it produced, with `win`, `killed_by`, `seed`, `deck_size`, `map_points`; 3. the A/B attribution row — `policy \t run \t code-hash \t session` — and the report prints each arm's session ids. The game's own run file does **not** know our session id and the harness originally mapped only a policy to a filename, so without (2) and (3) nothing lined up and there were still no labels. - `jev` is `null` when the decision asked nothing (code paths, fallbacks), which is a fact about the decision rather than a gap. - `RecordingClient` in `run.py` captures the answers, so nothing in `brain.py` had to change. `JEV_TRACE` still works and is still worth setting, but it is now redundant for labelling. It is no longer the prerequisite. --- ## 4. Where our Jev usage stands (from `09`, plus this audit) Measured usage of the documented surface: | Capability | Used | |---|---| | `Choice` | **7** questions — `best_play`, `which_potion`, `target`, `next_node`, `best_option`, `best_relic`, `best_bundle` | | `Noul` | **6** question families — `use_potion`, `want_any`, `skip_all`, `good_card{i}`, `good_relic{i}`, `good_card{i}` (card_select), `worth_item{i}` | | `Score` | **0** — though `score()` is defined at `jev.py:169` | | Structured criteria (`entry()`) | **0** | | Contrastive Noul criteria (`true=`/`false=`) | **0** | | `focus=` / `inspect=` / `compare=` | **0** | | `gate_choice()` / `margin()` | **0** in `brain.py` (`gate()` is used; `margin()` only inside `jev.py`) | Every question is a bare one-line string. The docs permit that only when the question is *short and unambiguous*; ours are neither, e.g.: ```python noul("Is any relic in `relics` worth taking over skipping?") ``` `09` §2 measured structured criteria on card play: **same card chosen 6/6**, margin 0.425 → 0.473. A small effect — worth adopting where disambiguation costs us, not as a blanket rewrite. That measurement stands. --- ## 5. The next iteration, in dependency order Ordered by dependency, not by appeal. Each step states what it unblocks. ### Step 1 — Enable measurement (blocks everything else) **Partly landed 2026-09-22.** The answers and the join are in; the corpus is not. Done: - every decision row carries the model's questions and answers, or `null` when nothing was asked (`run.py:decision_record`, `RecordingClient`); - a unique `session` id on every row, and `capture/sessions.jsonl` at exit mapping it to the run file(s) produced, with the outcome fields; - `ab_card_skip.sh` stamps the same session into its attribution rows. Still to do: - capture **all** state types, not just combat (`run.py` currently dumps `live_{step}_combat.json` only); - namespace captures per run so they stop overwriting across sessions; - log the *features* the decisions turn on (the `features` block below is a sketch, not what is written today — the questions and answers are the evidence currently recorded). ```json {"run_id":"1790...","step":42,"state_type":"card_reward", "features":{"card0_power":2.1,"card0_synergy":1.4}, "action":"select_card_reward:0","source":"jev","confidence":0.71, "gate":"pass"} ``` **Why first:** it is the prerequisite for every quality question in `09` and for continual learning (§6). With the join in place, the first usable dataset is one A/B run away; what is missing now is coverage (all state types, namespaced per run) and volume, not mechanism. ### Step 2 — Re-run and check whether the ceiling moved The defense fix (`08`, failure #35) is the single highest-leverage change made so far. Verify it with the A/B harness, now that `--stop-on-run-end` makes one session equal one run: ```bash ./ab_card_skip.sh 6 # 6 runs per arm, one run per session ``` **The question:** does `THE_KIN_BOSS` stop killing 9 of 37? If yes, runs start reaching Act 2+ and the reward signal finally varies. If no, the defense fix is wrong and nothing downstream is worth building. **Why second:** it is the test of the only substantive finding, and it gates §6. ### Step 3 — Match primitives to answer shapes Only after Steps 1–2 give labels: - Card quality: **`Score` with explicit ordered levels**, per card, several atomic axes, combined with weights in code (composite scoring). Replaces the `Noul` + `argmax` that the docs say is the wrong shape. - Structured criteria on the fuzzy Nouls (`worth taking`, `stronger`). - Measure each change against the Step 1 trace before keeping it. ### Step 4 — Then deck composition Archetypes, synergy, removal priority. Last, because `08` shows the bot does not currently reach the point where it matters. --- ## 6. Continual learning — the design, and the honest blocker The docs give the mechanism directly: > "Combine independent answers with deterministic rules or weighted sums. **For > learned composition, use the probabilities as features in a downstream > classical machine-learning model.**" So continual learning here is **not fine-tuning Jev**. It is: Jev produces features, code combines them, and **the weights are learned from outcomes**. Four stages: 1. **Instrument** — Step 1 above. 2. **Label** — attach the run outcome to each decision row. 3. **Fit** — a small logistic regression or gradient-boosted tree on `features → P(reach Act 2)`. The coefficients *are* the composite weights. 4. **Close the loop** — retrain periodically and A/B the new weights against the old, using the `--stop-on-run-end` harness. The docs also prescribe how to set our gates, which we have so far guessed: > "Test thresholds by plotting confidence against accuracy on your data." `EVENT_TOP_MIN = 0.60`, `EVENT_MARGIN_MIN = 0.30` and `CONFIDENCE_FLOOR = 0.55` were chosen by hand. With a trace they can be **measured**. (Note that `CONFIDENCE_FLOOR` is inert where it is passed: `gate()` ignores its threshold argument for a `ChoiceAnswer`, so the target question is gated by the margin rule, not by 0.55.) ### The blocker, stated plainly **We have 0 wins in 37 runs.** The top of the reward scale is unobserved, so there is nothing to learn from yet. Our only varying signal is map points reached — noisy, and confounded by seed luck. Training on this data would fit the bug, not the game. **Step 2 is what creates the gradient.** If runs start reaching Act 2 and Act 3, a label appears (`P(reach Act 2)`) that actually varies, and Stage 3 becomes possible. --- ## 7. Refactor implications `brain.py` is ~1700 lines with 13+ handlers, and it now mixes four concerns: state parsing, arithmetic, question building, and policy. The three-rule framing in §1 suggests the split that the code keeps trying to make: ``` sts2bot/ game/ card and encounter knowledge, mechanics constants state/ observation parsing, the five index spaces, CombatFacts policy/ one module per state_type; chooses among enumerated actions model/ TypeSafe primitives, a shared question library, the trace learn/ decision log, labels, weight fitting eval/ A/B harness, metrics, run-file mining ``` Two rules that should hold after the split: 1. **`policy/` may not do arithmetic.** If a decision needs a number, the number comes from `state/`. This is the rule the defense bug broke. 2. **`model/` may not be asked a question whose answer is arithmetic.** This is the rule the `Noul`-for-spectrum bug broke. Both are testable as structural assertions, in the same style as the existing `test_brain.py` hostile-index tests (indices from data, `OFFSET = 5`). **Do not refactor yet.** The measurement in Step 1 should land first: it is cheap, it is a prerequisite for everything, and moving files while the corpus format is still changing would mean doing it twice. --- ## 8. Open questions carried forward | Question | Blocked on | |---|---| | Do the gates (0.45 / 0.60 / 0.30) match their accuracy? | **per-decision labels** — replay, expert judgement, or a code-verifiable ground truth. Run outcomes cannot answer this: one run result attached to a step does not say whether that step's answer was correct. | | Is structured criteria worth adopting beyond card play? | trace, then A/B | | Does the two-stage cascade beat one broad question? | an optimal-action label | | Does map beam search beat the single-step Choice? | `map_point_history` labels | | Is `TREMBLE` (offered 45×, taken 0×) correctly rejected, or is that a bug? | a card-quality label | | Does the event keyword veto reject benign options? | event captures + outcomes | | Is the escalation tier (act / review / escalate) needed? | trace | --- ## 9. Summary - **Three rules**, one per evidence source, and every serious bug broke one. - **One fix landed** (defense) with the highest measured leverage: 9 runs, one boss, one cause. - **One blocker**: 0 wins means no reward gradient, so continual learning is designed but cannot start. - **One prerequisite**: enable `JEV_TRACE` and log decisions. Cheap, and it unblocks every quality question above. - **One test that matters next**: re-run and see whether `THE_KIN_BOSS` deaths drop. That single number decides whether the rest of this plan is worth building.