Join the three evidence sources that were kept apart, and extract the useful
result: each source independently decides what belongs in code and what belongs
to the model, and every serious bug we have found broke one of those rules.
* Game mechanism -> arithmetic and rules belong in CODE
* Internal state -> anything directly readable or writable belongs in CODE
* TypeSafe -> judgements go to the MODEL, in the primitive matching the
answer's shape
* never blocked in boss fights -> game mechanism (arithmetic left to the model)
* card quality asked as a Noul -> primitives (a spectrum forced into yes/no)
* indices read from the wrong array -> internal state
The counter-example is kept too: the Score we tried for fight-level planning was
unusable, so the rule is "match the primitive to the shape, and verify" rather
than "use Score more".
Also records the ordered plan (measurement first, then re-run, then primitives,
then deck composition), the continual-learning design, and why it is blocked:
with 0 wins in 37 runs the reward has no gradient, so training would fit the bug
rather than the game.
4.4 KiB
STS2 Bot — Research Notes
Living notes on reverse-engineering Slay the Spire 2 and driving it with a TypeSafe System One (Jev) decision model.
These documents record what was measured, not what was assumed. Where a claim comes from a live test, the evidence is quoted. Where something is a guess, it says so.
Contents
| # | Document | Covers |
|---|---|---|
| 01 | Game engine and mod surface | Engine, assemblies, the official mod loader, manifest schema |
| 02 | System One / Jev | What Jev is, measured latency and cost, the arithmetic failure |
| 03 | STS2MCP interface | The community mod, version drift, rebuilding from source |
| 04 | State shapes | Every verified JSON shape, per state_type |
| 05 | Failure modes | Every infinite loop found live, with its fix |
| 06 | Decision architecture | The three-layer design and why the split is where it is |
| 07 | Run log | Results per run, with seeds and outcomes |
| 08 | What actually wins | Why the web meta is unusable, and the measured cause of 9 run losses |
| 09 | TypeSafe best practice | Doc-by-doc gap analysis: what we match, what measurement killed, what is missing |
| 10 | Synthesis and next iteration | Game + state + primitives joined; the three rules; the planned iteration and refactor |
Architecture summary lives in ../DESIGN.md.
How to add a finding
- Prefer a measurement over an inference. Quote the command and the output.
- Put game facts in 01/03/04, model facts in 02, bugs in 05.
- Record the date and the game build (
v0.107.1today). Both move. - When a finding is later disproved, do not delete it. Mark it superseded and say what replaced it. The wrong turn is often the useful part.
- Re-measure numbers before you copy them. Test counts, latencies and run totals in these docs have gone stale more than once. The suites print their own totals; use those, not a figure quoted from an earlier session.
Worked example: a hand-off summary recorded
29and118assertions. Re-running the suites showed50and131, and the difference was two regression sections that already existed in the files. The stale numbers would have gone into this index unchallenged. Run the suite; do not trust the summary.
Environment these notes were taken on
| Item | Value |
|---|---|
| Game build | v0.107.1, commit 59260271 |
| Platform | macOS (arm64), Steam |
| Engine | Godot 4.5.1 (.NET), runtime .NET 9.0.7 |
| Mod | STS2MCP, rebuilt from upstream main @ 55e0648 |
| Model | jev-latest resolving to jev-1.13.0 |
Rebuilding the mod
STS2MCP release 0.4.0 is broken on this game build. See
03. To rebuild:
cd ~/sts2-bot/vendor/STS2MCP
nix shell nixpkgs#dotnet-sdk_9 --command bash -c '
dotnet build STS2_MCP.csproj -c Release -o out/STS2_MCP \
-p:STS2GameDir="$HOME/Library/Application Support/Steam/steamapps/common/Slay the Spire 2"'
cp out/STS2_MCP/STS2_MCP.dll \
"$HOME/Library/Application Support/Steam/steamapps/common/Slay the Spire 2/SlayTheSpire2.app/Contents/MacOS/mods/"
Then restart the game. Mods load only at process start.
Testing
Three suites, all offline. Run them before every session.
python3 test_brain.py # 131 assertions — structural / programmatic
python3 test_facts.py # 50 assertions — arithmetic and parsing
python3 test_run.py # 21 assertions — the decision log and its joins
test_brain.py is the important one for catching usage bugs. It asserts that
every state_type produces an action legal for that state, that every
action the decision layer can emit is declared somewhere, and that every
fallback respects its own inputs. It needs no model and no running game.
test_run.py drives main() end to end with a fake sts2 and a stub client,
so it pins the decision log without a network: a model decision's row carries
that decision's answers, and a decision that asked nothing logs jev: null
rather than inheriting the previous step's. Removing the reset before each
decide() makes it fail.
It was added after a session in which four separate infinite loops and three hardcoded fallbacks were found by hand. Most of them would have been caught here.