From baf0d529d3015f19e7e7980e78f77bdaf9c87688 Mon Sep 17 00:00:00 2001 From: 0xrsydn Date: Wed, 23 Sep 2026 14:17:38 +0700 Subject: [PATCH] docs(policy): document composite scoring boundaries and limitations --- docs/COMPOSITE_REWARDS.md | 70 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 70 insertions(+) create mode 100644 docs/COMPOSITE_REWARDS.md diff --git a/docs/COMPOSITE_REWARDS.md b/docs/COMPOSITE_REWARDS.md new file mode 100644 index 0000000..74aa154 --- /dev/null +++ b/docs/COMPOSITE_REWARDS.md @@ -0,0 +1,70 @@ +# Experimental composite reward scoring + +Enable with `python3 run.py --card-skip-policy composite --stop-on-run-end`. +This command makes model calls and executes game actions. Run it only when authorized. +The default remains `jev`. Shops, combat, and other selection screens are unchanged. + +`policy/reward_scoring.py` builds context from observed reward metadata, player +resources, relics, and scoped combat-pile evidence. It preserves cost, rarity, +keywords, upgrades, and descriptions. It does not reconstruct the permanent deck. + +One model call asks four independent Score questions per offered card: + +| Dimension | Weight | +| --- | ---: | +| First-shuffle offense | 0.35 | +| Direct defense | 0.30 | +| Long-fight scaling | 0.15 | +| Draw, resource, and setup consistency | 0.20 | + +Each dimension has five descriptive levels. Code normalizes each fractional score +with `(score - 2) / 2`, then computes the weighted sum. Rarity is context, not a +bonus. Pairing matters through observed support, not a separate synergy bonus. +The weights and positive-addition margin of 0.10 are experimental, not calibrated. +Utility is not a win probability or confidence. + +## Module boundaries + +- `policy/scoring.py` contains pure normalization, weighted utility, and stable + ranking functions. It has no Jev dependency, game rules, or acceptance thresholds. +- `policy/reward_scoring.py` owns reward context, rubrics, model calls, answer gates, + weights, and reward acceptance/fallback rules. +- `policy/selection.py` dispatches the selected reward policy. +- `run.py` executes decisions and records the scoring audit. + +Shared arithmetic rejects missing components, invalid numbers, and invalid weights. +Weights must be nonnegative and sum to one. No component is silently removed. +Normalization maps each side of the neutral point linearly onto [-1, 1]. Ranking +preserves input order for ties. Callers retain responsibility for evidence quality. +Other policies can reuse these functions without sharing reward prompts or weights. +This extraction does not change reward behavior or resolve the skip limitation. + +## Acceptance and limits + +- Missing scoped card evidence or no model client uses the existing rarity fallback. +- Missing, nonfinite, out-of-range, or low-confidence components use that fallback. +- All four components for every card must pass. The current Score confidence gate + is 0.60. This is a heuristic gate, not proof of correct reasoning. +- Valid results select the highest weighted utility. Ties preserve offer order. +- A forced selection takes the highest utility even when all utilities are negative. +- **Automatic composite skipping is not implemented.** The serializer exposes + `can_skip`, but does not establish the alternative's effects. A weak or harmful + set therefore selects the highest-ranked card as a labeled fallback. This can + still add an unwanted card. Do not treat this version as a complete take/skip policy. + +The permanent deck, card instance identity, and skip effects remain integration +prerequisites. Combat evidence can contain generated cards and temporary upgrades. +The model must not assume missing support or future upgrades. Better context does +not establish better decisions; compare recorded decisions before promoting this mode. + +Proposal and compatibility records contain `scoring`: policy version, weights, +margin, evidence provenance, raw components, acceptance flags, utilities when +available, and the actual gate reason. Existing linked model records retain the +questions and parsed responses. No outcome becomes a per-decision correctness label. + +Offline regression coverage is in `test_brain.py`. It checks batched questions, +fractional arithmetic, indices, context preservation, fallback paths, forced +selection, and recording. No live game or model connection is needed. + +See [the research proposal](research/15-composite-card-policy-proposal.md) for the +full algorithm and evaluation plan. This implementation is its first, limited stage.