From 62693618db17df596ab9bc97c041d7052f06505f Mon Sep 17 00:00:00 2001 From: 0xrsydn Date: Tue, 22 Sep 2026 12:45:46 +0700 Subject: [PATCH] docs(dev): add lean agent guidance and prototype hardening research --- .pi/README.md | 11 + .pi/prompts/sts2-check.md | 16 + .pi/prompts/sts2-review.md | 15 + AGENTS.md | 44 +++ CONTEXT.md | 40 +++ docs/research/11-prototype-hardening.md | 392 ++++++++++++++++++++++++ 6 files changed, 518 insertions(+) create mode 100644 .pi/README.md create mode 100644 .pi/prompts/sts2-check.md create mode 100644 .pi/prompts/sts2-review.md create mode 100644 AGENTS.md create mode 100644 CONTEXT.md create mode 100644 docs/research/11-prototype-hardening.md diff --git a/.pi/README.md b/.pi/README.md new file mode 100644 index 0000000..4ef5ea1 --- /dev/null +++ b/.pi/README.md @@ -0,0 +1,11 @@ +# Project Pi prompts + +Start Pi from the repository root and trust this project to load its prompts. +Use `/reload` after adding or editing prompts in an already trusted session. + +- `/sts2-check`: run offline tests, dataset integrity checks, and shell syntax checks. +- `/sts2-review [focus]`: review changes for correctness and data integrity without editing files. + +`AGENTS.md` contains shared development instructions. +No project settings, packages, or extensions are required. Global model preferences remain unchanged. +Prompts guide the agent; they do not enforce permissions or isolate commands. diff --git a/.pi/prompts/sts2-check.md b/.pi/prompts/sts2-check.md new file mode 100644 index 0000000..03f03c5 --- /dev/null +++ b/.pi/prompts/sts2-check.md @@ -0,0 +1,16 @@ +--- +description: Run offline STS2 development checks and summarize failures +--- +Read AGENTS.md and inspect `jj status` before checking the current workspace. +Run these commands from the repository root: + +```sh +python3 test_facts.py && python3 test_brain.py && python3 test_run.py +python3 migrate.py --check-only +bash -n eval_batch.sh ab_card_skip.sh +``` + +Run the dataset and shell checks even if a Python test fails. +Do not run the game loop, call Jev, rebuild the dataset, or modify source files. +Report commands, exit codes, test totals, and any environment-dependent skips. +Finish with `jj status` and identify any unexpected workspace changes. diff --git a/.pi/prompts/sts2-review.md b/.pi/prompts/sts2-review.md new file mode 100644 index 0000000..3f4d9ed --- /dev/null +++ b/.pi/prompts/sts2-review.md @@ -0,0 +1,15 @@ +--- +description: Review STS2 changes for regressions and data integrity +argument-hint: "[focus]" +--- +Read AGENTS.md, then inspect `jj status` and `jj diff`. +Review the current change. If the change is empty, review the requested area of the current code. +Additional focus: ${ARGUMENTS:-game-loop safety and dataset integrity} + +Check fresh observations after actions, card-index validity, deterministic arithmetic, +model-free fallbacks, confidence gates, transport failures, and dry-run side effects. +Check session/step attribution, state/action linkage, and separation of run outcomes from correctness labels. +Use offline reproductions where useful. Do not call live services or modify files. + +Report actionable findings first, ordered by severity, with file/line references and missing regression tests. +Separate confirmed defects from proposed improvements. State review limits and checks performed. diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..4ec89e3 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,44 @@ +# Development + +## Code map +- `facts.py`: pure game-state parsing and arithmetic. +- `brain.py`: decision policy; returns one `Decision` per observation. +- `run.py`: observe–decide–act loop, captures, and session attribution. +- `sts2.py`: local game HTTP client. `jev.py`: TypeSafe model client and gates. +- `migrate.py`: dataset migration and integrity checks. +- `CONTEXT.md`: domain terms. `docs/DATASET.md`: dataset schema and limits. +- `docs/research/11-prototype-hardening.md`: audited findings and the proposed development order. +- Treat historical claims in `docs/DESIGN.md` as context; verify against current code. + +## Checks +Use Python 3.10+ from the repository root. Core Python code uses the standard library. + +```sh +python3 test_facts.py && python3 test_brain.py && python3 test_run.py +python3 migrate.py --check-only +bash -n eval_batch.sh ab_card_skip.sh +``` + +Tests are executable scripts, not unittest/pytest suites. They need no model or game connection. +Some checks read local captures and game history when available. +Prefer integration/end-to-end checks; keep only essential regression tests and use temporary isolated probes for unit-level edge cases. + +## Invariants +- Execute one game action, then read a fresh observation. Card indices can change after each action. +- Compute arithmetic and legality in code, not in Jev. Confidence does not prove correctness. +- Keep deterministic fallbacks usable with `client=None`. Test model paths with stubs. +- Preserve session/step attribution. Run outcomes are not per-decision correctness labels. +- Split evaluation data by run, not by decision row. + +## Safety and scope +- Default to offline tests. Ask before live game actions, model calls, or evaluation batches. +- `run.py --dry-run` sends no action POSTs or deck-cache writes; state reads, model calls, and logs still run. +- Never print or commit credentials. Do not read secret files for development checks. +- Do not rebuild `dataset/`, change captures, or edit `vendor/` unless the task requires it. +- `migrate.py --verify` rebuilds data; use `--check-only` for routine verification. +- Keep code simple and lean. Use `docs/` as the detailed reference for humans and agents. +- Keep changes focused. Prefer `jj status` and `jj diff`; do not create commits unless asked. + +## Pi +Project prompts live in `.pi/prompts/`: `/sts2-check` and `/sts2-review`. +They provide instructions, not a sandbox or permission enforcement. diff --git a/CONTEXT.md b/CONTEXT.md new file mode 100644 index 0000000..a7e5d19 --- /dev/null +++ b/CONTEXT.md @@ -0,0 +1,40 @@ +# STS2 Bot + +Terms for observing Slay the Spire 2, choosing actions, and evaluating play. + +## Language + +**Run**: +One game attempt, from its start to victory, death, or abandonment. +A seed describes generated content; it does not uniquely identify an attempt. + +**Session**: +One period of bot operation. A run can span sessions, and a session can span runs. + +**Observation**: +The game information visible at one moment. It is not the complete game state. + +**Fact**: +A directly observed value or a result established from known game rules and sufficient inputs. +_Avoid_: Using this term for an unverified prediction. + +**Estimate**: +A prediction that depends on assumptions or incomplete game information. +_Avoid_: Guaranteed result, proven lethal + +**Decision**: +A proposed next action. A decision does not establish that the game received or applied the action. + +**Action result**: +The game's response to an attempted action. Acceptance alone does not establish that the expected change occurred. + +**Run deck**: +The persistent collection of cards owned during a run, distinct from temporary combat cards and combat piles. +_Avoid_: Treating all visible combat cards as the persistent deck. + +**Run outcome**: +The result and progress of one game attempt. It does not establish whether each earlier decision was correct. + +**Decision label**: +An assessment of one decision against a stated criterion, such as legality or an expert preference. +_Avoid_: Treating the run outcome as an optimal-action label. diff --git a/docs/research/11-prototype-hardening.md b/docs/research/11-prototype-hardening.md new file mode 100644 index 0000000..a536322 --- /dev/null +++ b/docs/research/11-prototype-hardening.md @@ -0,0 +1,392 @@ +# 11 — Harden the prototype before reorganizing it + +Reviewed: 2026-09-22. +Status: original audit and recommendations, not a policy-quality benchmark. + +Implementation update: the first correctness pass fixes several findings below. +See [12 — Correctness pass](12-correctness-pass.md) for current behavior, checks, and remaining work. + +## Recommendation + +Keep Python, synchronous execution, and the existing facts/policy/client separation. +Make the prototype safe to test and measure before moving it into a large package structure. + +The problem is not that the code is simple. The problem is that some estimates are treated as facts, +and proposed actions can change policy memory before the game applies them. +A directory refactor would preserve these defects unless tests expose them first. + +Use three small work packages: + +1. Correct mechanics assumptions and action failure handling. +2. Capture complete decision evidence and make policy memory explicit. +3. Add a focused test harness, then extract modules when an actual change needs the boundary. + +Do not wait for a winning run before improving testability. Do not treat better organization as evidence of better play. + +## Evidence and limits + +This review used the current Python code, vendored STS2MCP source, the checked-in dataset, +existing research, and current public TypeSafe documentation. +The upstream STS2MCP API reference fetched during this review matches the vendored reference byte for byte. +The upstream `main` commit returned by GitHub was `55e064850a68f3b4cde7e5fd525bf9b2dec4e885`. +This does not verify which binary a running game has loaded. + +No game actions, paid Jev requests, secret reads, dataset rebuilds, or game-save changes were made. + +Offline checks: + +- Existing scripts: **50 + 131 + 21 assertions passed**. +- Dataset integrity: **37 runs, 1,052 decisions, 346 unique observations**. +- Independent snapshot replay: **346 observations, no policy exceptions with `client=None`**. +- Combat coverage: **335 observations**, but only **93** have both a player play phase and enemies. +- The corpus has **no `boss`, `event`, or `shop` observations**. +- Outcomes: **0 wins; 30 runs with one act entered and 7 with two**, using the dataset's `progress.acts_entered` field. + +Successful replay is not proof of good actions. Observations are not ordered trajectories. +The audit resets screen guards between snapshots, so it cannot establish correct transition behavior. +The script tests proposal behavior separately with synthetic observations. + +Reproduce the audit from the repository root: + +```sh +python3 utils/audit_prototype.py +python3 test_facts.py && python3 test_brain.py && python3 test_run.py +python3 migrate.py --check-only +bash -n eval_batch.sh ab_card_skip.sh +``` + +`audit_prototype.py` is a diagnostic, not a pass/fail gate. Exit zero means the audit completed. +Its output reports current behavior, including defects. Convert each confirmed defect into a regression test when fixing it. + +## Confirmed defects and misleading contracts + +### 1. Hand damage is not necessarily base damage + +`facts.py:damage_to_target`, `facts.py:_subset_damage`, and `brain.py:_lethal_line` +parse damage from card text and add Strength. + +The corpus contains two observations where: + +- Player Strength is 2. +- An unupgraded Strike in hand says `Deal 8 damage.` +- `damage_to_target(..., target_status=[], target_block=0)` returns **10**. + +Example: `dataset/states/06/0618279b1a19a34c32469d654d00386fa904e76fc1cf8ee0943d47469d257a49.json.gz`. + +The source explains the distinction. `McpMod.Helpers.cs:SafeGetCardDescription` calls +`GetDescriptionForPile(pile)`. `McpMod.StateBuilder.cs:BuildCardState` uses the hand description, +while pile entries use their respective pile descriptions. These are display values, not a stable base-damage API. +Captured Frail states likewise show Defend as `Gain 3 Block.`, not its unmodified value. + +**Next:** define a versioned observation contract for displayed damage versus base damage. +Use captured fixtures to establish which modifiers are already applied. Do not simply remove every modifier: +target-specific effects, conditional damage, multi-hit rounding, and damage-reducing powers still need verification. +Centralize the calculation so three separate functions cannot drift. + +### 2. Individual killability is not a whole-combat lethal proof + +`facts.py:combat_facts` tests each enemy with the entire hand and energy budget. +It sets `lethal_available` when every enemy is individually killable. + +Synthetic counterexample: + +- One Strike deals 6 single-target damage and costs 1 energy. +- The player has 1 energy. +- Two enemies each have 6 HP. +- Current result: both enemies are `killable`, and `lethal_available=True`. +- Actual result under these rules: only one enemy can die. + +`brain.py` currently executes a line against one enemy, so this does not mean it attempts both actions without observing. +It does mean the semantic fact sent to Jev is false. + +**Next:** keep `killable_targets` separate from a joint lethal result. +A joint result must respect shared cards, energy, targets, and known effects. +Until that exists, report whole-combat lethal as unknown rather than claiming a proof. + +The module's existing “lower bound” claim is also unsafe when damage-reducing enemy powers are ignored. +Ignoring a reduction can overestimate damage. + +### 3. The defense rule can spend energy on unnecessary block + +`CombatFacts.block_urgent` can remain true after all incoming damage is blocked. +`brain.py:combat_decision` then forces the largest block card before consulting any other policy. + +A synthetic case with 12 block, 12 incoming, a plain Defend, and a Strike still chooses Defend. +There is no block-retention power or other benefit in that fixture. + +One corpus snapshot also triggers this rule with 12 block and 11 incoming: +`dataset/states/2e/2ec44b9da79079a39b5f29b784ce1430b7e43f45866a2c5f455fb31430a5d636.json.gz`. +Its chosen card is Shrug It Off, which also draws a card. This demonstrates the rule firing, +not that this particular captured action is necessarily bad. + +`max_block_available` is also a greedy estimate, not a maximum. +With 2 energy, a 2-cost/9-block card beats two 1-cost/6-block cards in the greedy ordering. +It reports **9**, although **12** is reachable. + +**Next:** distinguish current block deficit from a fight-duration estimate. +Use exact small-hand optimization for supported effects, or name the value as an estimate. +Keep block retention, draw, and other side effects separate from immediate block value. + +### 4. A proposal changes selection memory before execution + +`brain.py:_card_select_pick` appends an index to `_card_select_chosen` while deciding. +Calling `decide` twice on the same unconfirmed card-selection observation, without sending any action, +produces `select_card` followed by `__wait__`: “1/1 chosen; waiting for confirm”. + +The same failure can occur when the game rejects the first selection. +Dry-run and offline evaluation also become stateful simulations of actions that never happened. +Similar proposal-time bookkeeping exists in shop and crystal-sphere handlers. + +**Next:** introduce explicit policy memory and a pending-action state. +Keep proposal generation free from changes that assert execution. +Record attempts separately, then reconcile the action result with the next observation. +An HTTP `ok` response is not sufficient evidence that an asynchronous change has completed. + +This is a worthwhile small refactor now. It is more important than splitting every handler into a file. + +### 5. Combat errors bypass the runner's model-failure budget + +`brain.py:combat_decision` catches `JevError` and returns a heuristic decision. +`run.py` sees success and resets its consecutive model-error counter. +An injected combat outage therefore does not reach the runner's `--max-jev-errors` handling. + +**Next:** give the runner one consistent failure signal. +Either let the error reach the runner or return structured model-attempt metadata with the fallback decision. +Separate unavailable-model failures from low-confidence answers. + +Also fix the previously identified dry-run writes and success exit codes on failed sessions. +Test dry-run at the actuator boundary, not only at the ordinary action branch. +Document that mod GET handlers can auto-open some screens; a no-POST mode is not a general game-state sandbox. +An offline observation file is the reliable no-game-side-effects preview. + +### 6. Answer logging does not record the actual policy gate + +`jev.py:answer_record` writes `gated` using the default gate. +Events apply stricter thresholds and a keyword veto inside `brain.py:event_decision`. + +For probabilities `{a: 0.55, b: 0.25, c: 0.20}`, the log says `gated=True`, +but the event's numeric gate rejects the answer. +That prevents accurate analysis of why the policy used or rejected a model answer. + +`JevClient._parse` also turns a missing `noul` field into `0.0`, which passes the Noul certainty gate. +Missing protocol data becomes a confident “no”. + +**Next:** validate question IDs, answer types, required fields, finite values, ranges, and candidate membership. +Log the gate name/version, thresholds, result, and veto reason where the policy makes the decision. +Keep the original answer distribution separate from that decision. + +### 7. Combat piles are not the persistent run deck + +`run.py` persists all-pile counts as the deck for macro decisions. +The audit finds **33 observations** where a hand Status card also appears in those deck counts. +One includes three Infection cards. Combat-generated cards, played powers, and temporary transformations +mean pile counts cannot establish the persistent deck without provenance. + +There is also no run identifier on `deck.json`; `STARTING_DECKS` exists but is unused. +A restarted process can initially load another run's snapshot. + +**Next:** label the current value as a combat-pile snapshot, with session/run identity and observation provenance. +Prefer a verified persistent-deck source. Otherwise maintain known deck changes and mark unknown information explicitly. +Do not silently present a combat snapshot as an exact run deck. + +## What Jev should do here + +The official architecture fits this project: ordinary code owns control flow, +while Jev supplies narrow semantic judgments. Keep that design. [1][2] + +| Primitive | Correct interpretation | Suitable STS2 use | Avoid | +|---|---|---|---| +| Noul | Probability that a specific proposition is true | Whether a card's described effect supports an explicitly described deck plan | Treating the value as damage, benefit magnitude, or win probability | +| Choice | Relative preference among supplied alternatives | Choose a card/target pair, event option, or reward including skip | Assuming an independent target answer knows which card another question selected | +| Score | Position on ordered, described levels | An experimental rubric for synergy or setup burden | Exact numerical predictions, or nominal categories such as Race/Trade/Control | + +### Keep batching, but couple dependent choices explicitly + +Questions in one request see the same state and are evaluated independently. [1] +Current combat asks `best_play` and one generic `target` question together. +A target that fits one card need not fit the selected card. + +Prefer either: + +- One Choice over valid `(card, target)` candidates, with computed immediate consequences; or +- One target question per candidate card in the same batch, then use the answer for the chosen card. + +A second request is needed only when the first answer supplies information that was not available earlier. +Add `end_turn` when it is a legitimate alternative. A card being playable does not establish that playing it helps. +Do not let a model invent executable parameters outside the candidate set. + +### Do not replace every Noul with Score + +The official re-ranking cookbook explicitly sorts candidates by Noul probability. [3] +The existing research is too categorical when it calls Noul plus argmax inherently wrong. + +The distinction is the objective: + +- “Which candidate most probably satisfies this clearly defined condition?” can use Noul ranking. +- “How much strategic benefit does this card provide?” needs a defined benefit rubric or a different evaluation design. + +A high probability of a small benefit is not the same as a large expected benefit. +Test alternatives on labeled examples; do not choose by confidence alone. +The previous Race/Trade/Control Score trial is not a general verdict against Score: +those are strategy categories, not clearly ordered levels of one quantity. [4] + +### Preserve uncertainty and version the experiment + +Structured criteria are useful for ambiguous boundaries, but strings are valid and often sufficient. [5] +Do not introduce structured wrappers everywhere merely because the API supports them. + +Choice confidence summarizes the distribution; it is not a direct probability that the action wins. [6] +The current margin gate is a defensible custom rule, not a universal requirement from TypeSafe. +Thresholds need domain labels and separate treatment for high-cost mistakes. + +Pin the model for experiments. The fetched model documentation maps `jev-latest` to `jev-1.13.0` +and warns that aliases can move. Record both requested and returned model IDs. [7] +Also record question, policy, feature, and game/mod versions. + +Jev's documented numerical limitations support keeping arithmetic in code. [2] +They do not make every Python calculation correct. Display semantics and incomplete mechanics still require validation. + +## Use the STS2 interface more precisely + +The vendored mod is the useful contract, not generic game guides. [8] + +- Card, reward, grid-selection, shop, and potion indices are observation-local handles. + Preserve their index spaces and validate against the observation that produced the candidate. +- Draw-pile entries are sorted for display in `BuildPlayerState`; their array order is not draw order. +- Hand and pile card descriptions have different display contexts. Preserve that provenance. +- Combat exposes information the policy does not fully support: stars, orbs, pets, and many powers. + State an initial support target, such as Ironclad singleplayer, rather than implying complete character support. +- `GET /api/v1/compendium` already exposes `current_run.run_id` and seed when the save is available. + `sts2.py` already wraps this endpoint. Use it at run boundaries instead of relying only on history-directory differences. + Handle absent or stale saved context, and strip `save_path` from published records. +- Enemy `combat_id` and `entity_id` are different identifiers. Use the current advertised `entity_id` for actions; + do not assume its generated suffix is a durable identity across observations. + +Code should validate an action against current observable constraints, then treat the game response as authoritative. +The static `LEGAL_ACTIONS` table is incomplete for cross-screen potion actions; do not promote it into a complete legality oracle. +Do not blindly retry a side-effecting POST after an ambiguous transport failure. Re-observe before deciding whether another attempt is valid. + +## The minimum useful architecture + +Keep the current modules until extraction pays for itself. Add only the boundaries needed for the next work: + +```text +raw observation + -> normalized view + supported/unknown mechanics + -> facts and explicitly named estimates + -> valid action candidates + -> deterministic rules + optional Jev judgments + -> decision proposal + -> runner attempts one action + -> action result + next observation reconcile policy memory +``` + +Small dataclasses or typed dictionaries are sufficient. Start with: + +- `PolicyContext`: session/run identity, deck knowledge, screen memory, and pending action. +- `Decision`: action proposal plus evidence and model-attempt metadata. +- `ActionAttempt`: attempt identity, observation reference, result, and later reconciliation. + +Preserve raw observations. Normalize only fields that the current policy consumes. +Unknown cost, missing intent, and unsupported effects must not silently become zero or “safe”. +Share mechanics calculations between facts and policy; avoid a blanket rule that policy cannot contain any arithmetic. + +A future package can expose `observations`, `mechanics`, `policy`, `clients`, and `recording` modules. +There is no need yet for services, queues, a plugin system, an asynchronous rewrite, or a learning subsystem. + +## Development order and acceptance checks + +### Work package A — correctness and failure visibility + +- Add captured fixtures for displayed damage and synthetic tests for shared energy and overblocking. +- Fix the mechanics contracts, not only the symptom in one duplicated function. +- Cover rejected selections, unchanged observations, transport failure, and every dry-run exit path. +- Make model failures visible to the runner and return explicit session stop reasons. + +Acceptance: new tests fail before each fix and pass afterward. Existing regression scripts still pass. +No claim of increased win rate follows yet. + +### Work package B — evidence and explicit state + +- Capture every observation before deciding, not only combat. +- Use content-addressed blobs or session-scoped paths. Link each decision to its exact observation. +- Record suppressed proposals, dry-run proposals, waits, attempts, accepted/rejected results, and session termination separately. +- Include run/session/step/attempt IDs, policy/model versions, actual gate decisions, and requested/returned model IDs. +- Extract explicit policy memory so interleaved sessions and offline comparisons cannot affect each other. +- Keep multi-call model evidence if later experiments introduce a cascade; `RecordingClient.last` would lose earlier calls. + +Acceptance: every attempted action has one exact pre-action observation and one recorded result or explicit unknown outcome. +Two sessions do not overwrite observations or share selection memory. +Interrupted runs remain visible rather than disappearing from evaluation. + +### Work package C — a lean development harness + +Testing policy, updated after user review: prefer integration and end-to-end checks. +Keep only essential permanent regressions. Use temporary isolated probes for low-level edge cases. + +- Run the real loop and clients against controlled local endpoints, with temporary storage and no live credentials. +- Keep the existing direct test commands; add dependencies or a test framework only when necessary. +- Run checks in continuous integration without a game, credentials, or personal history directories. +- Replace local-history-dependent assertions with temporary fixtures. +- Use a Nix shell when dependency isolation helps. It does not isolate game saves or prevent network access. +- Extract combat mechanics or a policy handler only when integration checks establish its contract. + +The official TypeSafe Python SDK is now available. [9] +An adapter is worth evaluating when replacing retry and protocol-validation code. +Do not change it concurrently with policy semantics. In particular, SDK Score maps use integer keys, +while the current local answer class expects string keys. + +## Evaluation: what the existing data can and cannot support + +**Zero wins is not zero learning signal.** The checked-in data already varies in act reach, +fight damage, turns, and progress. It cannot train a useful win classifier from positive examples it does not have. +It can support diagnosis, expert labeling, and carefully scoped intermediate metrics. +A tiny, policy-biased dataset still does not justify a claim of learned action value. + +Keep three evaluation levels separate: + +1. **Contract checks:** valid candidate references, supported mechanics, fresh observations, protocol handling, and state transitions. +2. **Labeled decision checks:** expert preferences or code-verifiable outcomes on a held-out, varied fixture set. +3. **Live policy checks:** act reach, deaths, fight damage/turns, intervention rate, failures, latency, and cost per completed run. + +Snapshot replay can detect regressions and compare proposals. It cannot reveal what would have happened after an unplayed action. +The current corpus does not provide sufficient state/action trajectories or behavior propensities for reliable counterfactual policy evaluation. +Run outcomes also cannot calibrate each Jev answer's correctness. + +For future live experiments: + +- Keep one policy fixed for an entire run, including resumed sessions. +- Interleave or randomize policy assignment instead of collecting all of one policy first. +- Record seed, character, ascension, game/mod build, unlock context, and termination status. +- Compare matching seeds only when the game mode actually supports controlled seeds. +- Split labels and evaluation by run, not by row. Do not use post-decision fight results as pre-decision features. +- Report small-sample uncertainty. Six runs per arm are a smoke test, not a conclusive strategy comparison. + +The next useful policy experiment is combat correctness and defense behavior, not another broad card-skip sweep. +Card drafting can matter in Act 1; the data does not prove otherwise. +The priority follows the concrete combat defects found here, not a claim that deck composition only matters later. + +## Changes to earlier conclusions + +These corrections take precedence over conflicting statements in research notes 08–10 and `docs/DATASET.md`: + +- Python arithmetic is not automatically an exact fact; display values can already include modifiers. +- Noul ranking is supported; the proposition and intended utility must be distinguished. +- The dataset already contains varied progress, despite zero wins. +- Run outcomes are not per-decision correctness labels. +- Snapshot replay is not counterfactual evaluation of policy quality. +- Current run identity and seed can be available through compendium, although absent from the basic live `run` object. +- `collect.py` is referenced in older documentation but does not exist in this codebase. +- A defense change being implemented is not evidence that the boss-survival problem is fixed. + +## Sources + +1. [TypeSafe primitives](https://docs.typesafe.ai/primitives) and [How to build with TypeSafe](https://docs.typesafe.ai/concepts/how-to-build-with-system-one). +2. [Jev 1.13 jaggedness](https://docs.typesafe.ai/model-jaggedness/jev-1.13). +3. [Re-ranking cookbook](https://docs.typesafe.ai/cookbooks/rerank_typesafe). +4. [Score](https://docs.typesafe.ai/primitives/score). +5. [Structured instructions and criteria](https://docs.typesafe.ai/primitives/advanced). +6. [Confidence](https://docs.typesafe.ai/confidence), [Noul](https://docs.typesafe.ai/primitives/noul), and [Choice](https://docs.typesafe.ai/primitives/choice). +7. [Models and version aliases](https://docs.typesafe.ai/models). +8. [STS2MCP API at the inspected revision](https://github.com/Gennadiyev/STS2MCP/blob/55e064850a68f3b4cde7e5fd525bf9b2dec4e885/docs/raw-simplified.md), plus local `McpMod.Helpers.cs`, `McpMod.StateBuilder.cs`, and `McpMod.Compendium.cs` under `vendor/STS2MCP/`. +9. [Official TypeSafe Python SDK](https://docs.typesafe.ai/sdk/python).