docs(research): synthesis of game, state and primitives; next iteration
Join the three evidence sources that were kept apart, and extract the useful
result: each source independently decides what belongs in code and what belongs
to the model, and every serious bug we have found broke one of those rules.
* Game mechanism -> arithmetic and rules belong in CODE
* Internal state -> anything directly readable or writable belongs in CODE
* TypeSafe -> judgements go to the MODEL, in the primitive matching the
answer's shape
* never blocked in boss fights -> game mechanism (arithmetic left to the model)
* card quality asked as a Noul -> primitives (a spectrum forced into yes/no)
* indices read from the wrong array -> internal state
The counter-example is kept too: the Score we tried for fight-level planning was
unusable, so the rule is "match the primitive to the shape, and verify" rather
than "use Score more".
Also records the ordered plan (measurement first, then re-run, then primitives,
then deck composition), the continual-learning design, and why it is blocked:
with 0 wins in 37 runs the reward has no gradient, so training would fit the bug
rather than the game.
This commit is contained in:
parent
f249349dd8
commit
65fe1171cd
2 changed files with 357 additions and 3 deletions
335
docs/research/10-synthesis-and-next-iteration.md
Normal file
335
docs/research/10-synthesis-and-next-iteration.md
Normal file
|
|
@ -0,0 +1,335 @@
|
||||||
|
# 10 — Synthesis: game, state and primitives → next iteration
|
||||||
|
|
||||||
|
Status: **forward-looking**. This doc is the join of three evidence sources that
|
||||||
|
until now were kept apart:
|
||||||
|
|
||||||
|
| Source | Doc | What it gives us |
|
||||||
|
|---|---|---|
|
||||||
|
| STS2 game mechanism | `01`, `08` | What actually decides a run |
|
||||||
|
| Internal game state we control | `04`, `08` | What is observable, and what we can change |
|
||||||
|
| TypeSafe / System One primitives | `02`, `09` | What shape to ask the model |
|
||||||
|
|
||||||
|
It records the design conclusions and the planned iteration. Nothing here is a
|
||||||
|
measurement unless it says so; the measurements live in `05`, `07`, `08`, `09`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. The synthesis in one line
|
||||||
|
|
||||||
|
**Each source independently decides what belongs in code and what belongs to the
|
||||||
|
model — and every serious bug we have found is a violation of one of those three
|
||||||
|
rules.**
|
||||||
|
|
||||||
|
| Source | Rule it imposes |
|
||||||
|
|---|---|
|
||||||
|
| Game mechanism | Anything that is arithmetic or a rule belongs in **code**. HP, block, lethal, damage projection, energy. |
|
||||||
|
| Internal state | Anything we can read or write directly belongs in **code**. Index spaces, save files, prefs, action enumerations. |
|
||||||
|
| TypeSafe primitives | Anything that is a *judgement over unstructured meaning* goes to the **model** — and must use the primitive that matches the answer's shape. |
|
||||||
|
|
||||||
|
### The three bugs, mapped to the three rules
|
||||||
|
|
||||||
|
| Bug | Rule broken | Doc |
|
||||||
|
|---|---|---|
|
||||||
|
| Bot never blocked in boss fights; 9 runs lost to `THE_KIN_BOSS` | **Game mechanism.** Blocking is arithmetic (`projected_incoming > affordable_loss`) and we left it to a damage-biased model question. | `08` |
|
||||||
|
| Card quality asked as a `Noul` and `argmax`'d | **Primitives.** A spectrum judgment forced into yes/no. `primitives.md`: "A Noul value of 0.5 means the model gives yes and no equal probability. It does not mean the candidate has a medium skill level." | `09`, §2 |
|
||||||
|
| Card indices read from the wrong array (5 spaces, similar names) | **Internal state.** We misread what the index pointed at. | `04` |
|
||||||
|
|
||||||
|
That mapping is the useful part. It turns three unrelated bug hunts into one
|
||||||
|
checkable rule, and it predicts where the next bug is: **any place we ask the
|
||||||
|
model something that is really arithmetic, or use the wrong primitive for the
|
||||||
|
answer shape.**
|
||||||
|
|
||||||
|
### The counter-example that keeps us honest
|
||||||
|
|
||||||
|
`09` §3 records a `Score` we tried for fight-level planning ("Race / Trade /
|
||||||
|
Control") that was **unusable on 4 of 6 states**, confidence as low as 0.01.
|
||||||
|
So the rule is not "use Score more". It is "match the primitive to the answer
|
||||||
|
shape, and verify the shape works before adopting it."
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. What the game actually rewards (from `08`)
|
||||||
|
|
||||||
|
Established by mining 37 run files, not by reading guides:
|
||||||
|
|
||||||
|
- **81% of runs (30/37) die in Act 1.** Deck composition is downstream of that.
|
||||||
|
- **43% die to an Act 1 boss; `THE_KIN_BOSS` alone kills 9.**
|
||||||
|
- The bot is **not** bloating, descriptively: 23% of rewards were skipped (249 of
|
||||||
|
323 rewards took a card), one card per reward at most. The per-card rate
|
||||||
|
(249/1087 = 23%) is mechanically diluted by offer size and does not diagnose
|
||||||
|
skip policy; we have no STS2 benchmark for a good skip rate.
|
||||||
|
- The Kin fights lasted **5–10 turns** and cost **44–80 HP**, ~10–13 a turn,
|
||||||
|
with block cards in hand. That is a defense failure, and it is now fixed.
|
||||||
|
|
||||||
|
**Consequence for planning:** any work on archetypes, synergy scoring or deck
|
||||||
|
"lean weight" is tuning a system the bot does not survive long enough to use.
|
||||||
|
The order is: survive Act 1 → then optimise the deck.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. What the state gives us that we are not yet using
|
||||||
|
|
||||||
|
Three levers found but not yet exploited:
|
||||||
|
|
||||||
|
### 3.1 The run files are a labelled dataset
|
||||||
|
|
||||||
|
`history/*.run` records, per map point: every card offered with `was_picked`,
|
||||||
|
`damage_taken`, `potion_used`, `rest_site_choices`, `upgraded_cards`, and the
|
||||||
|
final deck. That is a supervised dataset of our own drafting and combat, and it
|
||||||
|
already invalidated one hypothesis (bloat) and found one bug (defense).
|
||||||
|
|
||||||
|
**Not yet used for:** labelling decisions, or measuring the gates.
|
||||||
|
|
||||||
|
### 3.2 Instant Mode is a one-line lever
|
||||||
|
|
||||||
|
`prefs.save` holds `"fast_mode"` as plain JSON, and the mod's Instant Mode
|
||||||
|
checkbox is only `PrefsSave.FastMode = FastModeType.Instant`. Setting it to
|
||||||
|
`"instant"` removes animations and should roughly halve wall-clock per run.
|
||||||
|
**Needs a game restart** — prefs load at startup.
|
||||||
|
|
||||||
|
This matters because the binding cost of every experiment is wall-clock. At
|
||||||
|
~10 minutes per run, an A/B of 6v6 is two hours.
|
||||||
|
|
||||||
|
### 3.3 Model answers are now logged with the decision (not in a side trace)
|
||||||
|
|
||||||
|
**Corrected 2026-09-22.** This section previously said
|
||||||
|
`JEV_TRACE=capture/jev_trace.jsonl` was "the entire change". It is not, and it
|
||||||
|
was never sufficient on its own: `jev.py`'s trace records a `HH:MM:SS` stamp,
|
||||||
|
the questions and the answers — **no run, session, step, state or action**. Two
|
||||||
|
calls in the same second are indistinguishable, so the file cannot be joined to
|
||||||
|
a decision or to an outcome, which is the only thing a label needs.
|
||||||
|
|
||||||
|
What the decision log carries now, per decided action
|
||||||
|
(`run.py:decision_record`):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"session":"20260922T034419","step":55,"state_type":"monster",
|
||||||
|
"run":{"act":1,"floor":5},"action":"play_card","source":"jev",
|
||||||
|
"params":{"card_index":1,"target":"SHRINKER_BEETLE_0"},
|
||||||
|
"confidence":0.85,"reason":"jev chose Dismantle","error":null,
|
||||||
|
"jev":{"model":"jev-1.13.0","latency_s":0.71,
|
||||||
|
"questions":{...},"answers":{
|
||||||
|
"best_play":{"kind":"choice","choice":"card1","margin":0.79,"gated":true,
|
||||||
|
"probabilities":{"card0":0.01,"card1":0.89,...}},
|
||||||
|
"use_potion":{"kind":"noul","noul":0.41,"yes":false},
|
||||||
|
"which_potion":{"kind":"choice","choice":"potion0","margin":1.0,"gated":true}}}}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `session` is a unique id per process (`<timestamp>-<6 hex>`). It is the join
|
||||||
|
key only because three places now carry it:
|
||||||
|
1. every decision row in `capture/decisions.jsonl`;
|
||||||
|
2. `capture/sessions.jsonl`, written by `run.py` at exit (via `atexit`, so it
|
||||||
|
covers every exit path) and mapping the session to the history run file(s)
|
||||||
|
it produced, with `win`, `killed_by`, `seed`, `deck_size`, `map_points`;
|
||||||
|
3. the A/B attribution row — `policy \t run \t code-hash \t session` — and the
|
||||||
|
report prints each arm's session ids.
|
||||||
|
The game's own run file does **not** know our session id and the harness
|
||||||
|
originally mapped only a policy to a filename, so without (2) and (3) nothing
|
||||||
|
lined up and there were still no labels.
|
||||||
|
- `jev` is `null` when the decision asked nothing (code paths, fallbacks), which
|
||||||
|
is a fact about the decision rather than a gap.
|
||||||
|
- `RecordingClient` in `run.py` captures the answers, so nothing in `brain.py`
|
||||||
|
had to change.
|
||||||
|
|
||||||
|
`JEV_TRACE` still works and is still worth setting, but it is now redundant for
|
||||||
|
labelling. It is no longer the prerequisite.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Where our Jev usage stands (from `09`, plus this audit)
|
||||||
|
|
||||||
|
Measured usage of the documented surface:
|
||||||
|
|
||||||
|
| Capability | Used |
|
||||||
|
|---|---|
|
||||||
|
| `Choice` | **7** questions — `best_play`, `which_potion`, `target`, `next_node`, `best_option`, `best_relic`, `best_bundle` |
|
||||||
|
| `Noul` | **6** question families — `use_potion`, `want_any`, `skip_all`, `good_card{i}`, `good_relic{i}`, `good_card{i}` (card_select), `worth_item{i}` |
|
||||||
|
| `Score` | **0** — though `score()` is defined at `jev.py:169` |
|
||||||
|
| Structured criteria (`entry()`) | **0** |
|
||||||
|
| Contrastive Noul criteria (`true=`/`false=`) | **0** |
|
||||||
|
| `focus=` / `inspect=` / `compare=` | **0** |
|
||||||
|
| `gate_choice()` / `margin()` | **0** in `brain.py` (`gate()` is used; `margin()` only inside `jev.py`) |
|
||||||
|
|
||||||
|
Every question is a bare one-line string. The docs permit that only when the
|
||||||
|
question is *short and unambiguous*; ours are neither, e.g.:
|
||||||
|
|
||||||
|
```python
|
||||||
|
noul("Is any relic in `relics` worth taking over skipping?")
|
||||||
|
```
|
||||||
|
|
||||||
|
`09` §2 measured structured criteria on card play: **same card chosen 6/6**,
|
||||||
|
margin 0.425 → 0.473. A small effect — worth adopting where disambiguation
|
||||||
|
costs us, not as a blanket rewrite. That measurement stands.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. The next iteration, in dependency order
|
||||||
|
|
||||||
|
Ordered by dependency, not by appeal. Each step states what it unblocks.
|
||||||
|
|
||||||
|
### Step 1 — Enable measurement (blocks everything else)
|
||||||
|
|
||||||
|
**Partly landed 2026-09-22.** The answers and the join are in; the corpus is not.
|
||||||
|
|
||||||
|
Done:
|
||||||
|
|
||||||
|
- every decision row carries the model's questions and answers, or `null` when
|
||||||
|
nothing was asked (`run.py:decision_record`, `RecordingClient`);
|
||||||
|
- a unique `session` id on every row, and `capture/sessions.jsonl` at exit
|
||||||
|
mapping it to the run file(s) produced, with the outcome fields;
|
||||||
|
- `ab_card_skip.sh` stamps the same session into its attribution rows.
|
||||||
|
|
||||||
|
Still to do:
|
||||||
|
|
||||||
|
- capture **all** state types, not just combat (`run.py` currently dumps
|
||||||
|
`live_{step}_combat.json` only);
|
||||||
|
- namespace captures per run so they stop overwriting across sessions;
|
||||||
|
- log the *features* the decisions turn on (the `features` block below is a
|
||||||
|
sketch, not what is written today — the questions and answers are the
|
||||||
|
evidence currently recorded).
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"run_id":"1790...","step":42,"state_type":"card_reward",
|
||||||
|
"features":{"card0_power":2.1,"card0_synergy":1.4},
|
||||||
|
"action":"select_card_reward:0","source":"jev","confidence":0.71,
|
||||||
|
"gate":"pass"}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why first:** it is the prerequisite for every quality question in `09` and for
|
||||||
|
continual learning (§6). With the join in place, the first usable dataset is one
|
||||||
|
A/B run away; what is missing now is coverage (all state types, namespaced per
|
||||||
|
run) and volume, not mechanism.
|
||||||
|
|
||||||
|
### Step 2 — Re-run and check whether the ceiling moved
|
||||||
|
|
||||||
|
The defense fix (`08`, failure #35) is the single highest-leverage change made
|
||||||
|
so far. Verify it with the A/B harness, now that `--stop-on-run-end` makes one
|
||||||
|
session equal one run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./ab_card_skip.sh 6 # 6 runs per arm, one run per session
|
||||||
|
```
|
||||||
|
|
||||||
|
**The question:** does `THE_KIN_BOSS` stop killing 9 of 37? If yes, runs start
|
||||||
|
reaching Act 2+ and the reward signal finally varies. If no, the defense fix is
|
||||||
|
wrong and nothing downstream is worth building.
|
||||||
|
|
||||||
|
**Why second:** it is the test of the only substantive finding, and it gates §6.
|
||||||
|
|
||||||
|
### Step 3 — Match primitives to answer shapes
|
||||||
|
|
||||||
|
Only after Steps 1–2 give labels:
|
||||||
|
|
||||||
|
- Card quality: **`Score` with explicit ordered levels**, per card, several
|
||||||
|
atomic axes, combined with weights in code (composite scoring). Replaces the
|
||||||
|
`Noul` + `argmax` that the docs say is the wrong shape.
|
||||||
|
- Structured criteria on the fuzzy Nouls (`worth taking`, `stronger`).
|
||||||
|
- Measure each change against the Step 1 trace before keeping it.
|
||||||
|
|
||||||
|
### Step 4 — Then deck composition
|
||||||
|
|
||||||
|
Archetypes, synergy, removal priority. Last, because `08` shows the bot does not
|
||||||
|
currently reach the point where it matters.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Continual learning — the design, and the honest blocker
|
||||||
|
|
||||||
|
The docs give the mechanism directly:
|
||||||
|
|
||||||
|
> "Combine independent answers with deterministic rules or weighted sums. **For
|
||||||
|
> learned composition, use the probabilities as features in a downstream
|
||||||
|
> classical machine-learning model.**"
|
||||||
|
|
||||||
|
So continual learning here is **not fine-tuning Jev**. It is: Jev produces
|
||||||
|
features, code combines them, and **the weights are learned from outcomes**.
|
||||||
|
|
||||||
|
Four stages:
|
||||||
|
|
||||||
|
1. **Instrument** — Step 1 above.
|
||||||
|
2. **Label** — attach the run outcome to each decision row.
|
||||||
|
3. **Fit** — a small logistic regression or gradient-boosted tree on
|
||||||
|
`features → P(reach Act 2)`. The coefficients *are* the composite weights.
|
||||||
|
4. **Close the loop** — retrain periodically and A/B the new weights against the
|
||||||
|
old, using the `--stop-on-run-end` harness.
|
||||||
|
|
||||||
|
The docs also prescribe how to set our gates, which we have so far guessed:
|
||||||
|
|
||||||
|
> "Test thresholds by plotting confidence against accuracy on your data."
|
||||||
|
|
||||||
|
`EVENT_TOP_MIN = 0.60`, `EVENT_MARGIN_MIN = 0.30` and `CONFIDENCE_FLOOR = 0.55`
|
||||||
|
were chosen by hand. With a trace they can be **measured**. (Note that
|
||||||
|
`CONFIDENCE_FLOOR` is inert where it is passed: `gate()` ignores its threshold
|
||||||
|
argument for a `ChoiceAnswer`, so the target question is gated by the margin
|
||||||
|
rule, not by 0.55.)
|
||||||
|
|
||||||
|
### The blocker, stated plainly
|
||||||
|
|
||||||
|
**We have 0 wins in 37 runs.** The top of the reward scale is unobserved, so
|
||||||
|
there is nothing to learn from yet. Our only varying signal is map points
|
||||||
|
reached — noisy, and confounded by seed luck.
|
||||||
|
|
||||||
|
Training on this data would fit the bug, not the game. **Step 2 is what creates
|
||||||
|
the gradient.** If runs start reaching Act 2 and Act 3, a label appears
|
||||||
|
(`P(reach Act 2)`) that actually varies, and Stage 3 becomes possible.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Refactor implications
|
||||||
|
|
||||||
|
`brain.py` is ~1700 lines with 13+ handlers, and it now mixes four concerns:
|
||||||
|
state parsing, arithmetic, question building, and policy. The three-rule framing
|
||||||
|
in §1 suggests the split that the code keeps trying to make:
|
||||||
|
|
||||||
|
```
|
||||||
|
sts2bot/
|
||||||
|
game/ card and encounter knowledge, mechanics constants
|
||||||
|
state/ observation parsing, the five index spaces, CombatFacts
|
||||||
|
policy/ one module per state_type; chooses among enumerated actions
|
||||||
|
model/ TypeSafe primitives, a shared question library, the trace
|
||||||
|
learn/ decision log, labels, weight fitting
|
||||||
|
eval/ A/B harness, metrics, run-file mining
|
||||||
|
```
|
||||||
|
|
||||||
|
Two rules that should hold after the split:
|
||||||
|
|
||||||
|
1. **`policy/` may not do arithmetic.** If a decision needs a number, the number
|
||||||
|
comes from `state/`. This is the rule the defense bug broke.
|
||||||
|
2. **`model/` may not be asked a question whose answer is arithmetic.** This is
|
||||||
|
the rule the `Noul`-for-spectrum bug broke.
|
||||||
|
|
||||||
|
Both are testable as structural assertions, in the same style as the existing
|
||||||
|
`test_brain.py` hostile-index tests (indices from data, `OFFSET = 5`).
|
||||||
|
|
||||||
|
**Do not refactor yet.** The measurement in Step 1 should land first: it is
|
||||||
|
cheap, it is a prerequisite for everything, and moving files while the corpus
|
||||||
|
format is still changing would mean doing it twice.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Open questions carried forward
|
||||||
|
|
||||||
|
| Question | Blocked on |
|
||||||
|
|---|---|
|
||||||
|
| Do the gates (0.45 / 0.60 / 0.30) match their accuracy? | **per-decision labels** — replay, expert judgement, or a code-verifiable ground truth. Run outcomes cannot answer this: one run result attached to a step does not say whether that step's answer was correct. |
|
||||||
|
| Is structured criteria worth adopting beyond card play? | trace, then A/B |
|
||||||
|
| Does the two-stage cascade beat one broad question? | an optimal-action label |
|
||||||
|
| Does map beam search beat the single-step Choice? | `map_point_history` labels |
|
||||||
|
| Is `TREMBLE` (offered 45×, taken 0×) correctly rejected, or is that a bug? | a card-quality label |
|
||||||
|
| Does the event keyword veto reject benign options? | event captures + outcomes |
|
||||||
|
| Is the escalation tier (act / review / escalate) needed? | trace |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Summary
|
||||||
|
|
||||||
|
- **Three rules**, one per evidence source, and every serious bug broke one.
|
||||||
|
- **One fix landed** (defense) with the highest measured leverage: 9 runs, one
|
||||||
|
boss, one cause.
|
||||||
|
- **One blocker**: 0 wins means no reward gradient, so continual learning is
|
||||||
|
designed but cannot start.
|
||||||
|
- **One prerequisite**: enable `JEV_TRACE` and log decisions. Cheap, and it
|
||||||
|
unblocks every quality question above.
|
||||||
|
- **One test that matters next**: re-run and see whether `THE_KIN_BOSS` deaths
|
||||||
|
drop. That single number decides whether the rest of this plan is worth
|
||||||
|
building.
|
||||||
|
|
@ -18,6 +18,9 @@ guess, it says so.
|
||||||
| 05 | [Failure modes](05-failure-modes.md) | Every infinite loop found live, with its fix |
|
| 05 | [Failure modes](05-failure-modes.md) | Every infinite loop found live, with its fix |
|
||||||
| 06 | [Decision architecture](06-decision-architecture.md) | The three-layer design and why the split is where it is |
|
| 06 | [Decision architecture](06-decision-architecture.md) | The three-layer design and why the split is where it is |
|
||||||
| 07 | [Run log](07-run-log.md) | Results per run, with seeds and outcomes |
|
| 07 | [Run log](07-run-log.md) | Results per run, with seeds and outcomes |
|
||||||
|
| 08 | [What actually wins](08-what-actually-wins.md) | Why the web meta is unusable, and the measured cause of 9 run losses |
|
||||||
|
| 09 | [TypeSafe best practice](09-typesafe-best-practice.md) | Doc-by-doc gap analysis: what we match, what measurement killed, what is missing |
|
||||||
|
| 10 | [Synthesis and next iteration](10-synthesis-and-next-iteration.md) | Game + state + primitives joined; the three rules; the planned iteration and refactor |
|
||||||
|
|
||||||
Architecture summary lives in [`../DESIGN.md`](../DESIGN.md).
|
Architecture summary lives in [`../DESIGN.md`](../DESIGN.md).
|
||||||
|
|
||||||
|
|
@ -28,6 +31,15 @@ Architecture summary lives in [`../DESIGN.md`](../DESIGN.md).
|
||||||
3. Record the **date** and the **game build** (`v0.107.1` today). Both move.
|
3. Record the **date** and the **game build** (`v0.107.1` today). Both move.
|
||||||
4. When a finding is later disproved, do not delete it. Mark it superseded
|
4. When a finding is later disproved, do not delete it. Mark it superseded
|
||||||
and say what replaced it. The wrong turn is often the useful part.
|
and say what replaced it. The wrong turn is often the useful part.
|
||||||
|
5. **Re-measure numbers before you copy them.** Test counts, latencies and run
|
||||||
|
totals in these docs have gone stale more than once. The suites print their
|
||||||
|
own totals; use those, not a figure quoted from an earlier session.
|
||||||
|
|
||||||
|
> Worked example: a hand-off summary recorded `29` and `118` assertions.
|
||||||
|
> Re-running the suites showed `50` and `131`, and the difference was two
|
||||||
|
> regression sections that already existed in the files. The stale numbers
|
||||||
|
> would have gone into this index unchallenged. Run the suite; do not trust
|
||||||
|
> the summary.
|
||||||
|
|
||||||
## Environment these notes were taken on
|
## Environment these notes were taken on
|
||||||
|
|
||||||
|
|
@ -57,11 +69,12 @@ Then restart the game. Mods load only at process start.
|
||||||
|
|
||||||
## Testing
|
## Testing
|
||||||
|
|
||||||
Two suites, both offline. Run them before every session.
|
Three suites, all offline. Run them before every session.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python3 test_brain.py # 54 assertions — structural / programmatic
|
python3 test_brain.py # 131 assertions — structural / programmatic
|
||||||
python3 test_facts.py # 29 assertions — arithmetic and parsing
|
python3 test_facts.py # 50 assertions — arithmetic and parsing
|
||||||
|
python3 test_run.py # 21 assertions — the decision log and its joins
|
||||||
```
|
```
|
||||||
|
|
||||||
`test_brain.py` is the important one for catching usage bugs. It asserts that
|
`test_brain.py` is the important one for catching usage bugs. It asserts that
|
||||||
|
|
@ -69,6 +82,12 @@ every `state_type` produces an action **legal for that state**, that every
|
||||||
action the decision layer can emit is declared somewhere, and that every
|
action the decision layer can emit is declared somewhere, and that every
|
||||||
fallback respects its own inputs. It needs no model and no running game.
|
fallback respects its own inputs. It needs no model and no running game.
|
||||||
|
|
||||||
|
`test_run.py` drives `main()` end to end with a fake `sts2` and a stub client,
|
||||||
|
so it pins the decision log without a network: a model decision's row carries
|
||||||
|
that decision's answers, and a decision that asked nothing logs `jev: null`
|
||||||
|
rather than inheriting the previous step's. Removing the reset before each
|
||||||
|
`decide()` makes it fail.
|
||||||
|
|
||||||
It was added after a session in which four separate infinite loops and three
|
It was added after a session in which four separate infinite loops and three
|
||||||
hardcoded fallbacks were found by hand. Most of them would have been caught
|
hardcoded fallbacks were found by hand. Most of them would have been caught
|
||||||
here.
|
here.
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue