No description
  • Python 88.4%
  • C# 8.7%
  • Shell 2.9%
Find a file
0xrsydn d676bf067e fix(eval): end each A/B session with its run, and attribute runs to arms
Two defects made the earlier A/B numbers meaningless:

  * Without `--stop-on-run-end` a session contained several runs, so runs could
    not be attributed to an arm.
  * The script then read "the last 2N runs", which silently compared runs from
    earlier experiments rather than the two arms.

Pass `--stop-on-run-end`, and attribute each run record to its arm by diffing
the history directory around each session.

Also raise the step budget and warn on a thin sample: measured, a run still
going at step 600 produces NO run record, so a short budget biases the sample
toward runs that died early -- exactly the wrong bias for this question.
2026-09-22 06:09:16 +07:00
capture Add reference game-state captures 2026-09-22 00:02:24 +07:00
docs Add design doc and research notes 2026-09-22 00:01:22 +07:00
utils Migrate docdump tool from sts2-re 2026-09-22 00:01:37 +07:00
vendor Vendor STS2MCP mod source and 0.4.0 release DLL 2026-09-22 00:01:30 +07:00
.gitignore Ignore run artifacts and trace logs 2026-09-22 00:02:31 +07:00
ab_card_skip.sh fix(eval): end each A/B session with its run, and attribute runs to arms 2026-09-22 06:09:16 +07:00
brain.py fix(combat): block on projected fight damage, and three lethal-search defects 2026-09-22 06:09:16 +07:00
capture.py Add STS2MCP HTTP client and state capture tool 2026-09-22 00:01:44 +07:00
eval_batch.sh Add batch eval and card-skip A/B scripts 2026-09-22 00:02:18 +07:00
facts.py fix(combat): block on projected fight damage, and three lethal-search defects 2026-09-22 06:09:16 +07:00
jev.py feat(jev): structured question shapes, Score primitive, per-call answer record 2026-09-22 06:06:06 +07:00
run.py feat(run): log every decision with the answers and the run outcome 2026-09-22 06:06:06 +07:00
sts2.py Add STS2MCP HTTP client and state capture tool 2026-09-22 00:01:44 +07:00
test_brain.py fix(combat): block on projected fight damage, and three lethal-search defects 2026-09-22 06:09:16 +07:00
test_facts.py fix(combat): block on projected fight damage, and three lethal-search defects 2026-09-22 06:09:16 +07:00
test_run.py feat(run): log every decision with the answers and the run outcome 2026-09-22 06:06:06 +07:00