# STS2 decision dataset — schema and usage An open, educational dataset of Slay the Spire 2 decisions, collected by driving the game with a TypeSafe System One (Jev) policy and mining the game's own run history. > Slay the Spire 2 is © Mega Crit. Game-derived identifiers are included for > research and education. This project is not affiliated with Mega Crit and is > not commercial. ## Build it ```bash python3 migrate.py --verify # existing data -> dataset/ python3 collect.py # going forward, records live decisions ``` `migrate.py` is idempotent: it rebuilds `dataset/` from scratch every run. ## Layout ``` dataset/ manifest.json provenance, counts, limitations runs.jsonl one row per run decisions.jsonl one row per decision, with its outcome attached states_index.jsonl one row per unique observation states/ab/.json.gz content-addressed raw observations ``` States are **content-addressed**: the digest is the sha256 of the raw bytes. Measured on the first migration, **56% of captures were unique**, so 44% of storage is saved by deduplication and identical observations join for free. ## Current corpus | | | |---|---| | Runs | 37 | | Decisions | 1052 | | Unique states | 346 | | Storage | 0.42 MB (from 3.28 MB raw) | | Wins | **0** | Decisions by kind: | Kind | Rows | What it is | |---|---|---| | `card_reward` | 323 | Card offered after a fight; 1 of 3 picked, or skipped | | `potion_reward` | 190 | Potion offered | | `relic_reward` | 162 | Relic offered | | `event` | 120 | Event option chosen (raw shape, varies by event) | | `rest_site` | 97 | `SMITH` / `REST` / other | | `upgrade` | 73 | Which card was upgraded | | `ancient` | 43 | Ancient choice (raw) | | `shop_purchase` | 19 | Shop transaction | | `remove` | 10 | Card removed | | `transform` | 8 | Card transformed | | `enchant` | 7 | Card enchanted | ## Schemas ### `runs.jsonl` ```json {"schema":"sts2.run/1","run_id":"1789944034","seed":"...","character":"IRONCLAD", "ascension":0,"win":false,"killed_by":"VANTOM_BOSS","killed_by_event":null, "progress":{"acts_entered":1,"map_points":17}, "final":{"deck":[{"id":"STRIKE_IRONCLAD","count":4}],"deck_size":19, "relics":["BURNING_BLOOD"],"potions":[],"max_potion_slots":3}} ``` ### `decisions.jsonl` Every row carries its own outcome, so no join is required to train. ```json {"schema":"sts2.decision/1", "decision_id":"1789944034:act0:pt0:card_reward", "run_id":"1789944034","kind":"card_reward", "context":{"act":0,"map_point":0,"map_point_type":"monster","hp":74,"max_hp":80, "gold":109,"encounter":"NIBBITS_WEAK","room_type":"monster", "monsters":["NIBBIT"]}, "options":[{"id":"SETUP_STRIKE","picked":true,"floor_added":1}, {"id":"TREMBLE","picked":false,"floor_added":null}, {"id":"BLOOD_WALL","picked":false,"floor_added":null}], "chosen":"SETUP_STRIKE","skipped":false, "fight":{"damage_taken":12,"turns_taken":2,"hp_healed":6,"gold_gained":10}, "outcome":{"run_win":false,"run_map_points":17,"run_killed_by":"VANTOM_BOSS", "run_acts_entered":1}} ``` Fields worth knowing: - **`context`** is the observable state at that map point. It does **not** include the deck — the game does not record deck composition per map point. - **`fight`** is the dense reward signal: what the policy actually paid in HP and turns. `damage_taken` and `turns_taken` are the two numbers a combat policy is judged on. - **`outcome`** is the sparse, run-level signal. It is attached to every row. - **`options`** for reward kinds includes the rejected options, so this is a ranking dataset, not just a classification one. ### `states_index.jsonl` ```json {"digest":"sha256:abc...","path":"states/ab/abc....json.gz", "bytes_raw":5382,"bytes_gz":712,"state_type":"monster","menu_screen":null, "seen_in":["live_042_combat.json"]} ``` ## Known limitations These are recorded in `manifest.json` too, so the corpus states its own gaps. 1. **No per-step combat state/action pairs.** The session logs record the action and reason but not the observation, and captures exist only for combat. A combat step therefore has a state with no action, or an action with no state — never both. **They are deliberately not reconstructed.** `collect.py` records them going forward. 2. **No per-map-point deck composition.** The game records observed deltas (`cards_gained`, `cards_removed`, `upgraded_cards`) and the final deck. Folding forward from the known starting deck is left to the consumer, because doing it here would silently bake in an assumption. 3. **Zero wins.** The reward signal has no positive class. See "Training" below. 4. **Policy bias.** Every decision is what *this* policy did, which is not the same as what was correct. There is no optimal-action label. 5. **`event` and `ancient` rows are stored raw.** Their option shape varies by event and has not been normalised. ## Splitting **Split by `run_id`, never by row.** Two decisions from the same fight share almost all of their context; a row-level split leaks and reports a fake score. ```python import json runs = [json.loads(l) for l in open("dataset/runs.jsonl")] train_ids = {r["run_id"] for r in runs[:30]} rows = [json.loads(l) for l in open("dataset/decisions.jsonl")] train = [r for r in rows if r["run_id"] in train_ids] ``` ## Training — read this before using it The corpus supports three uses today, and one that is blocked. | Use | Status | |---|---| | Offline policy evaluation | ✅ replay states, score any policy, no game needed | | Regression corpus | ✅ freeze states, assert decisions do not change | | Gate calibration | ✅ confidence vs accuracy on real decisions | | Supervised fine-tuning | ⚠️ **imitates a policy that never won a run** | | Reinforcement learning | ❌ **no reward gradient — 0 wins** | **The trap:** this is a recording of a losing player. Distilling it into weights teaches the bot to reproduce the mistakes that lost 9 of 37 runs to a single Act 1 boss. The data is evidence about *what the policy did*, not about *what works*. For RL the blocker is worse than sample size. With `win=false` on every row the value function has nothing to climb — a reward that is 0 everywhere teaches that all actions are equally bad. **What unblocks it:** the corpus starts being useful for learning once runs reach Act 2 and Act 3, so the reward varies. That is a game-play problem (fixing why runs end early), not a data problem. The practical middle path, and the one the TypeSafe docs recommend, is to use Jev's probabilities as **features** in a small classical model rather than fine-tuning the model itself: > "For learned composition, use the probabilities as features in a downstream > classical machine-learning model." ## Integrity checks ```bash python3 migrate.py --verify # rebuild, then check python3 migrate.py --check-only # check an EXISTING dataset, no rebuild ``` **Use `--check-only` to verify.** `--verify` alone rebuilds first, so it can only ever see data that is correct by construction — it is a smoke test, not a check. This was found by trying to corrupt a row and watching `--verify` pass anyway. The checker asserts: - every state blob exists and **re-hashes to its recorded digest** - every `card_reward` with a pick has exactly one `picked: true` option - `chosen` is always one of `options` - no absolute paths and no API keys in the output - every row carries an `outcome` - every `run_id` in `decisions.jsonl` exists in `runs.jsonl` Each check was verified by deliberately breaking the dataset and confirming the checker fails. Measured, all five trip: | Injected fault | Result | |---|---| | `chosen` set to a value not in `options` | `FAIL ...: chosen not in options` | | 1 byte appended to a state blob | `FAIL hash mismatch states/2b/2bb595...json.gz` | | state blob deleted | `FAIL missing state states/1c/1cd4fc...json.gz` | | `/Users/...` written into `runs.jsonl` | `FAIL absolute path leaked into the corpus` | | `outcome` removed from a row | `FAIL ...: no outcome attached` | Exit code is `0` when clean and non-zero when not, so it works as a gate in CI before publishing.