# 09 — TypeSafe best practice: what we already do, and what we don't Status: **current**. Read against the vendor documentation on 2026-09-22: [How to build with System One](https://docs.typesafe.ai/concepts/how-to-build-with-system-one), [Advanced: structure](https://docs.typesafe.ai/primitives/advanced), [Choice](https://docs.typesafe.ai/primitives/choice), [Composite scoring](https://docs.typesafe.ai/patterns/composite-scoring), [Skill suggestion](https://docs.typesafe.ai/cookbooks/skill_suggestion), [Jev 1.13 jaggedness](https://docs.typesafe.ai/model-jaggedness/jev-1.13). Every claim below is either **measured** (numbers in this repo) or **by construction** (a synthetic case that reproduces the defect). Claims that need game-outcome evidence are labelled as such and are NOT implemented. --- ## 1. Where we already match the documentation | Documented practice | Where | |---|---| | Code owns control flow and side effects | `run.py` loop, `facts.py` arithmetic, `_lethal_line` | | Never ask the model to compare numbers | `lethal_available`, `survives_with_cards`, `killable`, buckets in `to_state()` | | Re-ranking: one absolute question per candidate, argmax in code | shop, card reward, card_select, relic_select | | Speculative fan-out: many questions, one call, code ignores what it needs | combat (4–5 questions per turn) | | Gate on margin, not confidence | `gate_choice`, `margin()` | | Count in code, one question per item | `good_card{i}` per offered card | | Backticked dot-and-index paths in questions | `combat.hand`, `offered.card{i}`, `items.{key}` | The bot's core shape — code enumerates the action space, Jev only judges between enumerated options — is exactly the architecture the guide describes. The gaps are in question *shaping*, not in the architecture. --- ## 2. Structured questions and criteria — available now, adopted nowhere The documentation allows an object or array wherever a string is accepted: `instructions`, every Choice option description, every Score level, and the Noul `criteria.true` / `criteria.false` entries. Until now the bot sent bare strings everywhere and never used `criteria` on a Noul at all. `jev.py` now builds those shapes: `entry()`, `ask()`, `noul(true=, false=)`, plus `gate_choice()` so each call site states its own Choice-calibrated floor. noul(ask("Would `offered.card0` make this deck stronger?", focus="Judge deck fit, not raw power.", inspect="`offered.card0`, `deck`"), true=entry("It adds damage, block or scaling this deck lacks"), false=entry("It is off-plan, redundant, or too slow to matter", not_for="A card that is merely different")) **Measured** (6 captured play-phase states, both shapes asked in the same call, so the state and the call are identical): the structured shape picked the **same card in 6/6**, mean margin 0.425 → 0.473, gate pass rate unchanged (5/6). That is a small effect on card play. It is worth adopting where disambiguation actually costs us — irreversible options (events), and the deck-fit questions whose absolute ratings are compressed — not as a blanket rewrite. --- ## 3. What the measurements killed ### A fight-level plan asked as a `Score` — unusable A three-level Score ("Race / Trade / Control") was asked on 6 real combat states: usable (confidence ≥ 0.5) on **2 of 6**; confidence as low as **0.01**, with `score` values straddling level boundaries (0.40–1.00). The jaggedness page warns that score levels are weak in numerical calibration, and it shows here. **Conclusion: do not route the fight-level view through Jev.** The code-side projection already in `facts.py` (`must_block`, `block_urgent`, `fight_is_grinding`, `this_turn_is_dangerous`) is the right tool, and it is arithmetic, which is exactly what belongs in code. ### "A bigger margin means a better play" — not supported by our data Stage two of the documented cascade (re-read the top three in detail, one Noul per candidate, free to reject all) was measured on the states where stage one came back flat: margins 0.02 → 0.28 and 0.16 → 0.73, with the per-candidate Nouls agreeing with the second Choice. But the same stage-one question, repeated on the same captured state in a later call, returned **0.04 → 0.24** and **0.36 → 0.02**. We have **no optimal-action label**, so a higher margin is not evidence of a better play — it is evidence that the margin is noisy near the boundary. **Conclusion: the cascade is a hypothesis, not a result.** It needs an outcome metric (HP lost per fight, or the run record) before it is wired in. It is deliberately not implemented. --- ## 4. Documented practice we are still missing ### 4.1 The event gate rejects the model on a third of event decisions `brain.py` sets `EVENT_TOP_MIN = 0.60` and `EVENT_MARGIN_MIN = 0.30` deliberately: events are irreversible, so the bar is higher than combat, and it was raised after a measured event option ended a run at confidence 0.49. Across the recorded trace that gate rejected Jev on **34 of 102** event decisions (20 "uncertain", 14 "risky"), leaving 17 where the model's answer stood. **Keep the gate as it is.** A conservative barrier on an irreversible decision is the right trade, and there is no event-labelled data — no outcome, no capture — that could justify relaxing it. The docs' warning about not carrying a Noul threshold onto a Choice says nothing about how these two constants were chosen, and lowering them to the combat rule (0.45 / 0.20) would weaken a measured safety barrier on a guess. The one thing worth revisiting later is the keyword veto, and only with data. It was built from one catastrophic case ("Lose Everything") and its word list contains ordinary English — `keep`, `more`, `continue`, `all of`, `again`, `deeper` — so it also vetoes options whose *title* is benign (`Trudge On`, `Solo Quest`, `Slowly Find an Exit`) on wording from the description. Narrowing it to irreversible wording (`everything`, `lose`, `sacrifice`, `all of`) is a candidate change, not a finding. **Blocker:** `capture/` holds no event states and the trace holds no event outcomes. Capture events, with the run result, first. ### 4.2 Decomposition of the broad combat question `best_play` asks one broad question ("which single play best advances winning this fight?"), which the guide calls the single most important thing to avoid. The documented alternatives are atomic per-axis questions composed with weights in code (composite scoring) or the two-stage cascade. Both need the outcome metric from §3 before adoption. ### 4.3 Beam search over the map graph The map state exposes the graph (`map.nodes[].children: [[col,row],…]`, `next_options[].leads_to`), and the hierarchical-classification cookbook prescribes exactly this shape: score candidates at each level and keep the best `K` paths instead of committing to one greedy step. Map pathing is currently a single-step Choice (mean confidence 0.57, the lowest of any decision, with 14 fallbacks). Path choice IS measurable — the run record carries `map_point_history` — so this is a good candidate for an outcome-scored A/B. ### 4.4 Escalation as a third tier The confidence-routing pattern has three tiers: act, review, escalate. We have two: act, or fall back to a deterministic heuristic. DESIGN.md promises escalation to a reasoning model for macro and it is not built. Given §3, the missing tier is where an independent signal belongs — not another margin. ### 4.5 State we hand the model - Enemy statuses used to be names without amounts, so Strength 6 and Strength 1 were indistinguishable. Now `{name, amount}` for both sides. - Macro decisions now receive `F.deck_context(deck)`: the full name→count map **plus** stable all-pile aggregates (size, kind mix, upgrades). The counts stay — card identities are the synergy and redundancy signal — and the aggregates are the part that is arithmetic. Every aggregate is stable across snapshots of the same deck, which is a real constraint, not a nicety: **there is no cost aggregate**, because the state does not expose a card's *base* cost and a temporary in-combat modifier applies to the copy in hand. Measured over 600 captures, one Strike read `cost: "0"` in hand while the same Strike read `cost: "1"` in the draw pile — an average cost would differ between two snapshots of an identical deck. `run.py` persists `{"counts", "summary"}` and `deck_context` still accepts a legacy flat snapshot, so an existing `deck.json` keeps working. Wired at all six macro call sites (card reward, map, card select, shop, treasure, bundle). Verified against the live model on `capture/10_card_reward.json`, and on a **two-option map fixture** built from `capture/06_map.json`: the capture itself offers a single next option, so `map_decision` short-circuits with `[code]` before any model call and proves nothing about the context. **The effect on decision quality is unmeasured** — this is additive context, not a proven improvement. --- ## 5. Correctness fixes landed with this doc All three are by construction, independent of any model question. | Fix | Before | After | |---|---|---| | Enemy block counted twice in lethal search (`facts._subset_damage`, `brain._lethal_line`) | hp 10 / block 5 vs 18 raw damage read as **not lethal** (required hp + 2·block) | lethal; block subtracted once | | The executor ignored the player's statuses (`_lethal_line(..., [], enemy)`) | `facts` said `lethal_available: true, enemies_you_can_kill_now: ['E_0']` while the code meant to execute the kill found nothing | `CombatFacts.player_status` is passed; the line is found | | `relic_select` asked `good_relicN` and read `relicN` | every answer missed, `best_by_noul` returned `(None, 0.0)` for every state, so the path could only ever take the rarest relic | ranks the prefixed ids and maps the winner back with the same prefix | `test_facts.py` **50/0** and `test_brain.py` **131/0** after all three, with regression cases for each: a blocked enemy, a Strength-carrying player, and a relic offer whose highest-rated relic is deliberately the *common* one so the rarity fallback cannot pass by accident. --- ## 6. Measurement prerequisites (these gate everything above) 1. **Model answers were discarded.** `JEV_TRACE` was never set, and it would not have been enough on its own: `jev.py`'s trace records a `HH:MM:SS` stamp and the answers, with **no session, step, state or action**, so two calls in one second are indistinguishable and nothing joins to a decision or an outcome. Fixed without touching `brain.py`: `run.py` wraps the client in `RecordingClient`, so every decision row in `capture/decisions.jsonl` carries the questions and the answers (`noul`, `probabilities`, `margin`, `gated`), or `null` when the decision asked nothing. Each row also carries a unique `session` id, and `run.py` writes `capture/sessions.jsonl` at exit mapping that session to the history run file(s) it produced, with the outcome fields. `ab_card_skip.sh` stamps the same id into its attribution rows. Join: `decisions.jsonl` (answers) → `session` → `sessions.jsonl` (outcome). `JEV_TRACE` still works and is now redundant. **What that supports, and what it does not.** It gives *arm-level* attribution — this decision belongs to this session, and the session's outcome is the run file — which is what an A/B or a delayed-reward model needs. It does **not** label individual answers: one run result attached to a step says nothing about whether that step's answer was correct, so "plot confidence against accuracy" cannot be answered from run outcomes. Gate calibration needs per-decision labels: replay, expert judgement, or a ground truth the code can verify on its own (lethal, legality, affordability). 2. **Only combat states are captured.** `run.py` dumps `live_{step}_combat.json` and nothing else, so shop, event, card_reward and card_select decisions cannot be measured offline. 3. **`live_{step}` collides across sessions** — each session overwrites the previous session's states, so the corpus cannot be segmented per run.