sts2-bot/docs/POLICY.md
0xrsydn 9696282110 feat(recording): link session observations proposals and action results
Capture all successful state reads with hashes and session-local IDs. Record action intent before POST, retain accepted/rejected/unknown results, and link subsequent observations. Preserve legacy feeds and finalize each invocation synchronously.
2026-09-22 15:53:14 +07:00

3.3 KiB

Policy structure

Current boundaries

run.py                 observe, execute, report results
recording.py           linked observations, proposals, attempts, and results
run_state.py           reported run identity and scoped combat-pile evidence
brain.py               entry point, dispatch, navigation, shops, minigames
policy/
    __init__.py        package marker; no registration or initialization
    context.py         Decision, PendingAction, PolicyContext
    combat.py          combat proposals and model questions
    selection.py       card, relic, bundle, hand, and reward proposals
facts.py               observation parsing and combat calculations

brain.decide() remains the policy entry point. The runner owns a RunContext, which holds policy memory and replaces it on reported run changes. The runner reports action results through that context. brain.Decision and brain.PolicyContext remain available as direct imports of the shared types, so the runner interface does not change.

Policy modules may ask Jev for preferences. They must not call the game API. context.py does not call either service. Combat arithmetic remains in facts.py, not duplicated inside combat policy.

The dependency direction is simple:

  • The runner uses brain.
  • brain uses the combat, selection, and context modules.
  • Combat and selection use context, facts, and the model client.
  • Context uses only the Python standard library.
  • No policy module imports brain or the runner.

State and configuration

A proposal does not update execution memory. Accepted requests are reconciled with fresh observations. See policy-state behavior and limits for the evidence required by each guarded action.

The existing card-skip setting stays in brain.CARD_SKIP_POLICY. The dispatcher passes its value explicitly to selection.card_reward_decision(skip_policy=...). This avoids a circular import and preserves the existing command-line behavior. It is configuration, not per-screen action memory.

Why stop here

These are plain modules, not a plugin framework or class hierarchy. Navigation, shops, and minigames stay together until a further split helps development.

The extraction does not change game decisions, confidence gates, fallbacks, state reconciliation, or recording. The subsequent run-state pass adds reported identity and card-evidence provenance. Linked recordings provide the corresponding session journal.

Extraction verification

  • 262 existing assertions passed. No test files or assertions were added for the extraction.
  • Thirteen isolated whole-process scenarios passed with local HTTP fixtures.
  • All 346 stored observations produced identical decisions before and after extraction, in both model-free mode and local-model-stub mode: 692 comparisons.
  • The local model requests also matched, including their state and questions.
  • Thirty-six function/class definitions retained identical abstract syntax trees. Only the two dispatch functions and explicit card-skip parameter changed.
  • Dataset integrity, Python compilation, and shell syntax checks passed.

The differential and whole-process probes were temporary. They made no live game or paid model calls. These checks establish refactoring equivalence for the tested inputs, not correctness for every possible game state.