DESIGN.md covers the three-layer architecture (facts in code, Jev for tactics, gated escalation for macro). research/ documents the engine and mod surface, the Jev classifier's measured behavior, the STS2MCP HTTP interface, state shapes, failure modes, decision architecture, and a run log of the first four sessions.
19 KiB
05 — Failure modes
Every entry below was observed live, caused a real infinite loop or a stalled run, and has a fix in the code. None are theoretical.
The general lesson: this interface is full of actions that report
{"status": "ok"} while doing nothing. Never trust ok. Trust the state.
1. claim_reward(index=0) forever
rewards.items[] is rebuilt and re-indexed from 0 after every claim.
Claiming index 0 repeatedly reclaims the same slot forever.
Observed: 65 consecutive rejected claim_reward(index=0) calls.
Fix: claim right-to-left, items[-1].
2. A potion reward with full slots silently vanishes
When all potion slots are full, claim_reward on a potion reward returns
ok and is silently dropped. The item never leaves items[], so the loop
never terminates.
Observed:
potions: 3 of max_potion_slots 3
claim_reward(1) -> ok | Claiming reward: potion (Energy Potion)
items now: [gold, potion] <- unchanged
claim_reward(0) -> ok | Claiming reward: gold (19)
items now: [potion] <- re-indexed
Fix: if the last item is a potion and len(potions) >= max_potion_slots,
discard the weakest potion first, then claim. Weakness comes from a small
POTION_VALUE table in brain.py.
3. select_card twice on one index freezes the screen
On grid screens select_card toggles. The second call deselects, so the
state never changes and the loop stalls.
Observed: select_card(index=0) twice, then 10 unchanged reads.
Fix: when preview_showing and can_confirm are true, the action is
confirm_selection, never another select_card.
4. A stale preview makes confirm_selection a no-op
confirm_selection returns ok while changing nothing, when the preview was
left over from a desynchronised state.
Observed:
confirm_selection -> ok "Confirming selection from preview"
state_type before/after: card_select -> card_select
sig changed: False
Fix: if the identical card_select state is seen twice in a row, the
confirm did not take effect. Send cancel_selection to reset, then re-select.
Verified working:
cancel_selection -> preview False
select_card(0) -> preview True
confirm_selection-> state_type becomes 'event'
5. rest_site options use name, not title
Reading options[].title yields None for every option, so the handler
concluded there were no options and fell through to a rejected proceed.
Fix: read name (plus id), and honour is_enabled.
6. treasure claims a relic before the chest opens
The chest auto-opens. During opening the response has no relics key and no
can_proceed, so claim_treasure_relic is rejected repeatedly.
Observed: 4 consecutive rejections.
Fix: wait while relics is absent; claim only when present.
7. Acting during transitions
The loop re-decided faster than the game animated:
choose_map_nodefired 3 times in a row during a single travel.end_turnfired repeatedly during the enemy turn.
Fix, two parts:
__wait__whenbattle.is_play_phaseis false.- An unchanged-state guard in
run.py: if the state signature repeats and the last action succeeded, wait. If the last action was rejected, re-decide instead, so a different action can be tried.
The second condition matters: waiting on a rejected action would prevent recovery and stall until the stuck counter fires.
8. unknown and overlay treated as dead ends
Both are transitions, not terminal states. Returning "stop" ended the session while the game was mid-transition into combat.
Fix: wait and re-observe. The unchanged-state guard bounds this, so a genuine dead end still stops after 10 reads.
9. The confidence gate rejected a correct answer
Not a loop, but a wrong decision, and worth recording because the bug was in our gate rather than in the model.
Measured: 5 cards offered. Jev picked Bash at 0.61 probability with
Defend at 0.29, and reported confidence 0.50. A fixed 0.55 floor rejected
it and fell back to a heuristic that chose to block instead.
confidence measures peakedness, so it falls as the option count rises:
(5 × 0.61 − 1) / 4 = 0.50.
Fix: gate a Choice on margin over the runner-up
(top >= 0.45 and top - runner >= 0.20), which is scale-free. See
02.
Checklist for a new state_type
- Capture the real shape. Do not code from documentation.
- Check which fields are absent in some states.
- Ask: what does this action return when it is a no-op?
- Prefer
__wait__over stopping for anything that looks transitional. - Add a guard so a repeated identical state does not loop forever.
Checklist for a new action
- Does it report
okon a no-op? (Most do.) - Does it re-index the collection it acts on?
- Is there a precondition the state exposes (
can_confirm,can_proceed,is_stocked,can_afford) that should be checked first? - Does it toggle?
Session 2 additions
10. An event option ended the run, and the gate allowed it
Measured on a deciphering event:
[102] jev chose Keep Deciphering conf=0.28
[104] jev chose Lose Everything conf=0.49
"Lose Everything" set the player's max HP to 1. The margin gate passed it, because the gate only measured how decisive the answer was, never how consequential the action was.
Fix, in two parts:
- A stricter gate for events:
top >= 0.60 AND margin >= 0.30. - A deterministic safety net. When the model is not confident, choose the
option with the lowest
event_safety_rank, which counts risk words ("everything", "keep", "continue", "gamble", "lose") minus stop words ("stop", "leave", "refuse", "decline", "take what").
A "does this risk losing the run?" Noul was tried and REMOVED. Measured on the same event:
| Option | Risk Noul |
|---|---|
| "Lose Everything" | 0.46 |
| "Keep Deciphering" | 0.52 |
| "Stop" | 0.38 |
It ranked the run-ending option as less risky than a moderate one. A misleading signal is worse than no signal, so danger is detected by keywords instead. Keep the model for the confident case; use code for the dangerous one.
11. hp=1/1 was reported as "healthy"
An effect reduced max HP to 1. _hp_bucket bucketed by percentage alone, so
1/1 was 100% and returned healthy. The bot walked into a normal fight at
1 HP and died.
Fix: absolute HP is now part of the bucket. hp <= 5 is always critical.
def _hp_bucket(pct, hp=None):
if hp is not None and hp <= 5:
return HP_CRITICAL
...
12. hand_select blindly selected index 0
combat_select_card(index=0) failed with:
Card index 0 out of range (0 selectable cards)
hand_select has cards (still selectable) and selected_cards (already
chosen). When cards is empty the only useful action is
combat_confirm_selection.
Fix: confirm when nothing remains selectable; otherwise give up basic Strikes first, then Defends.
13. bundle_select has the same preview trap as card_select
A bundle preview is already open - confirm or cancel it first
Same shape (preview_showing, can_confirm) and same fix: confirm when a
preview is showing, and reset with cancel_bundle_selection if the identical
state repeats.
14. Error text lives in error, not message
ActionResult read only message, so every rejection printed as
action rejected: with nothing after it. This hid four separate bugs for a
whole session.
Fix: message = data.get("message") or data.get("error") or "".
Lesson: make failures loud before chasing them. A blank error message is worse than no error handling.
15. can_proceed is not reliable
For shops, shop.can_proceed was false while proceed() worked and moved the
game to the map. Waiting on that flag stalls forever.
Fix: do not gate an exit on can_proceed. Attempt proceed and let the
rejection counter bound the retries.
16. The unchanged-state guard was count-based
A boss death animation plus the rewards transition exceeded 10 reads, so the guard declared STUCK while the game was still animating.
Fix: the guard is now time-based (--stuck-seconds, default 25 s).
17. Transient rejections are normal
proceed at a rest site right after a heal is rejected for a moment and then
succeeds. The retry loop was too impatient.
Fix: on rejection, back off 3x the normal pause, and allow 6 attempts.
Session 3 — the programmatic audit
Prompted by observing that the bot "was only upgrading common attack cards". These are usage bugs, not decision-quality issues, and are listed separately from the tuning items.
18. card_select fallback was hardcoded to cards[0]
The upgrade screen fell back to the first card in the list, and the list is ordered with basic Strikes first. So every fallback upgraded a Strike.
Measured across runs:
[156] [fallback] select_card(index=0) # select the first card
[263] [fallback] select_card(index=0) # low confidence; select the first
[378] [fallback] select_card(index=0) # low confidence; select the first
The fallback fired often because a single Choice over a 13+ card deck
dilutes, exactly like the shop.
Fix, two parts:
- Re-ranking: one absolute
Noulper card, argmax in code. - Screen-aware deterministic fallback (
upgrade_rank,removal_rank) that never picks index 0 blindly:- upgrade prefers a non-basic card, then Bash, then Strikes, then Defends, and never an already-upgraded card
- remove/transform invert the order: shed basic Strikes and Defends first, and never target an upgraded card
Verified live after the fix:
[197] jev chose Perfected Strike (noul=0.67)
[205] jev chose Bludgeon (noul=0.66)
[212] jev chose Rampage (noul=0.67)
[303] jev chose Bash (noul=0.62)
19. hand_select fed good cards to a "choose any number" prompt
The give-up ranking returned 2 for anything that was not a Strike or Defend,
so once the basics ran out it started offering real cards:
[023] give up Uppercut
[086] give up Stomp
[171] give up Bash <- the deck's only Vulnerable source
Fix: only basic Strikes and Defends are candidates. When none remain,
confirm and keep the good cards. A selection is forced only when
can_confirm is false.
20. card_reward had a redundant gate that skipped almost everything
A want_any Noul ("does this deck want any of these?") gated the whole
decision. When it was merely uncertain (0.54–0.59) the bot skipped, so it
skipped nearly every card reward and ran a 10-card deck.
Fix: the per-card Nouls ARE the signal. Skip only when no card clears
CARD_PICK_THRESHOLD.
21. A skipped card reward is NOT consumed — infinite loop
[409] skip_card_reward() # no card cleared 0.6
[410] claim_reward(index=2)
[411] skip_card_reward()
[412] claim_reward(index=2) ... forever
Skipping returns to the rewards screen with the card still listed. Claiming it again reopens the card screen, and the cycle repeats.
Verified directly: skip_card_reward -> ok, then the rewards list still
contains [2] card: Add a card to your deck.
Fix: record that a card reward was skipped, and ignore card rewards on the
rewards screen afterwards. rewards and card_reward are declared one
screen group so the flag survives the hop between them — otherwise it is
cleared on every transition and the loop returns.
22. Module-level guards leaked across screens
Found by the new test_brain.py. A fresh fake-merchant shop was reported as
"unchanged after a purchase" because a shop signature from an earlier screen
was still set.
Fix: _reset_screen_guards() clears all per-screen state whenever the
screen group changes. Also made the shop guard precise: it now only fires when
we actually purchased from that exact shop state.
23. fake_merchant nests its inventory one level deeper
The shop is at fake_merchant.shop.items, not shop.items. Reading only
obs["shop"] made every fake-merchant shop look empty, so the bot left
immediately without buying.
Fix: resolve obs["shop"], else obs["fake_merchant"]["shop"], else
obs["fake_merchant"].
24. embark is rejected without a character selected
action rejected: Embark button not available — select a character first
An earlier version assumed embark defaults to the first unlocked character.
It only worked once because a character happened to be selected already.
Fix: read the game's own message:
"Select a character."-> select one"Selected The Ironclad. Use 'confirm' to embark."-> embark
25. relic_select used a diluted Choice and ignored can_skip
Verified shape: relic_select.relics[] with index/id/name/description/rarity
plus can_skip. It now uses the same re-ranking pattern, honours can_skip,
and falls back to the rarest relic rather than index 0.
New: test_brain.py
A structural regression suite, deliberately separate from decision quality. It asserts:
- Every
state_typeproduces an action that is LEGAL for that state. Compares againststs2.LEGAL_ACTIONS. This is what would have caught a handler emittingend_turnduring someone else's turn. - Every action the decision layer can emit is declared somewhere, so a typo cannot silently produce an invalid action.
- Every fallback respects its inputs — removal does not target an
upgraded card, upgrade does not target a basic one, card reward takes the
rarest rather than index 0,
hand_selectgives up a Strike rather than Bash.
54 assertions, runs offline with client=None, no model calls, no game.
Run it before every session.
Session 4 — the toggle and transition traps
Found by running the bot for long stretches. Every one of these was a stall or a crash, not a decision-quality issue.
26. combat_select_card TOGGLES — re-selecting deselects
The same trap as card_select, in a different action. Measured: 18
consecutive combat_select_card(card_index=0) # give up Defend with the state
never changing, because selecting an already-selected card deselects it.
hand_select.selected_cards is the authoritative list of what is already
chosen. Exclude those from the candidate list.
Watch out: cards[].index and selected_cards[].index are different
index spaces (different arrays, independent counters). Names are the only
reliable way to match between them.
27. hand_select has a mode that only needs confirming
Measured state:
mode = "upgrade_select"
prompt = "Confirm Card to Upgrade"
cards = [(0, "Defend")]
selected = []
can_confirm = true
Sending combat_select_card is a silent no-op here. The card is already picked;
combat_confirm_selection closes the screen and moves to monster.
Fix: mode-aware.
| mode | action |
|---|---|
upgrade_select |
combat_confirm_selection |
simple_select |
select cards, then confirm when nothing selectable remains |
28. Character select: no indicator, and flaky embark
Two problems at once:
- The state carries no "selected" indicator. The mod hardcodes
result["message"] = "Select a character."whatever is chosen. Verified inAddCharacterSelectMenuState. So the message cannot be used to tell whether a character is picked. - Embarking right after selecting is flaky. Measured three consecutive
Embark button not available - select a character firstrejections, because the selection had not registered yet.
Fix: ALTERNATE — select, embark, select, embark. A rejected embark is always followed by a fresh select, so the sequence is self-correcting regardless of timing.
An earlier attempt used signature comparison ("if the screen is unchanged, a character was already selected"). It failed because the signature keeps changing during the embark transition, restarting the cycle.
29. The run loop's guard blocked the brain's own recovery
run.py waited whenever the state was unchanged after a successful action, and
it did so before calling brain.decide. That made it impossible for any
handler to notice a repeated state and react differently — which is exactly how
character_select and multi-line dialogue work.
Fix: move the check to AFTER deciding, and suppress only a repeated action, never acting in general. Proposing a different action is always allowed.
30. A bounded retry is required
Even the corrected guard stalled on Ancient dialogue: one click was issued, the state had not updated yet, and the guard then refused to retry forever.
Fix: --max-duplicate-waits (default 3). After that many suppressions on an
unchanged state, re-execute the action. Some actions legitimately need repeating
and some transitions are just slow.
31. unknown and overlay are transitions, not screens
Resetting the per-screen guards on them wiped state mid-flow. During embark the
state flickers through unknown, which cleared the character-select guard and
restarted the select/embark cycle.
Fix: _reset_screen_guards returns early for unknown and overlay.
32. menu_screen must be part of the screen key
main, singleplayer, character_select and tutorial_prompt all share
state_type == "menu" but have completely different valid actions. Keying only
on state_type meant moving from the main menu to character select did not
reset anything.
Fix: _screen_group() returns menu:<menu_screen> for the menu state.
33. RemoteDisconnected crashed a run
http.client.RemoteDisconnected: Remote end closed connection without response
RemoteDisconnected is an http.client.HTTPException, not a URLError, so
it escaped jev.py's transport handler and killed the process mid-fight.
Fix, both layers:
jev.pyalso catcheshttp.client.HTTPExceptionandOSError, and retries.run.pywrapsbrain.decidein a broadexcept Exception, falls back to the deterministic handlers, and keeps playing.
34. Known, self-correcting: stale-state rejections
Occasionally the game rejects a card we chose:
Card 'Infection' cannot be played: HasUnplayableKeyword
Card 'Evil Eye' cannot be played: EnergyCostTooHigh
can_play comes from the game's own card.CanPlay(), so the state was correct
when read. A Jev call takes ~0.75 s, so the state we decided on can be about a
second stale by the time the action lands.
This is a race, not a logic bug, and it self-corrects: the rejection sets
last_ok = False, so the loop re-decides immediately against a fresh state
instead of waiting. Cost is one wasted action. No revalidation pass is needed.
Test coverage after this session
test_brain.py: 78 assertions. New this session:
- indices must come from the data, never from list position (hostile-index
cases where
indexdeliberately differs from array position) hand_selectskips already-chosen cards, including duplicate nameshand_selectconfirms forupgrade_select- character select alternates select/embark and resets correctly
- a
StubClientmakes the model paths testable offline and deterministically