sts2-bot/docs/research/02-system-one-jev.md
0xrsydn fb32822468 Add design doc and research notes
DESIGN.md covers the three-layer architecture (facts in code, Jev for
tactics, gated escalation for macro). research/ documents the engine and
mod surface, the Jev classifier's measured behavior, the STS2MCP HTTP
interface, state shapes, failure modes, decision architecture, and a
run log of the first four sessions.
2026-09-22 00:01:22 +07:00

199 lines
7.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 02 — System One / Jev
Everything here is measured against `jev-latest``jev-1.13.0` from
`https://api.typesafe.ai/v1/systemone`.
## What Jev is
TypeSafe states it plainly:
> "System One is TypeSafe's model for building AI-powered software, not agents.
> It does not generate code or choose its own next action."
It is a **calibrated classifier**. It evaluates a `state` and answers typed
questions about it. It cannot plan, cannot call tools, and cannot invent values
outside the option set you supply.
| Primitive | Shape | Returns |
|---|---|---|
| `Choice` | pick 1 of ≤ 255 options | `choice`, `probabilities`, `confidence` |
| `Score` | position on ordered levels | `score`, `legend`, `probabilities`, `confidence` |
| `Noul` | yes/no | `noul` (probability only, **no confidence**) |
Documented properties: questions in one call are evaluated **independently and
in parallel**. One answer is never hidden context for another.
## Measured latency — 7× the documented figure
The docs say "most queries complete in about 100 ms". Measured end-to-end from
this machine:
| Request | Time |
|---|---|
| Minimal (1 short noul) | 0.73 s |
| Minimal × 5 repeats | 0.73, 0.73, 0.74, 0.80, 0.76 s |
| Full combat state, 3 questions | 0.90 s |
| `GET /v1/models` TTFB | 0.68 s (connect 0.22 s) |
The floor is **~0.73 s**, dominated by network RTT plus server TTFB.
**The useful consequence:** a 1-question call and a 3-question call over a full
combat state differ by only 0.17 s. State size and question count are nearly
free. **Batch every question for a state into one call.**
## Cost
| | |
|---|---|
| Input | $0.042 per Mtok |
| Output | free |
| Rate limit | 250k tok/s, 1,200 req/min |
| Context | 64k total; 32k for `state` + longest question |
| Input type | text only (string, JSON object, array) |
For comparison, the STS2MCP README reports a full run costs a frontier LLM
about **8M tokens** (~$2040). The same volume through Jev is roughly
**$0.080.35 per run**.
## THE CRITICAL FINDING: Jev cannot do arithmetic
Ground truth: `energy 3`, hand `Strike(1 cost, 6 dmg) ×3` plus
`Bash(2 cost, 8 dmg)`, target on `19 HP`. Maximum reachable damage is **18**,
so lethal is **NO**.
| Question | Jev answered | Verdict |
|---|---|---|
| Max damage bucket | `18_to_23` @ 0.73 | **Correct** |
| Is lethal available? | `0.79` | **Wrong** |
Jev bucketed the magnitude correctly but failed the threshold comparison of 18
versus 19 — **and reported 0.79 confidence on the wrong answer**.
Two conclusions:
1. Jev is decent at *approximate magnitude*, poor at *exact comparison*.
2. **Confidence gating cannot protect against arithmetic errors.** A
0.79-confidence wrong answer passes any sane threshold.
This is why `facts.py` computes every sum, comparison, and threshold, and hands
Jev only conclusions (`"lethal_available": true`). This is a regression test:
see `test_facts.py` case 1.
The jaggedness page confirms the general shape: Jev "does not count reliably",
"is not a calculator", degrades with indirection, and suffers "context rot"
when the state carries irrelevant detail.
## The confidence gate trap
**A fixed confidence floor is wrong when the option count varies.**
Measured: 5 cards offered. Jev picks `Bash` at **0.61**, with `Defend` at 0.29.
Reported `confidence` is **0.50**.
Why: confidence measures *peakedness*. For 5 options the documented formula is
`(count × peak 1) / (count 1)` = `(5 × 0.61 1) / 4` = 0.50.
A 0.55 floor rejected a clear plurality and fell back to a worse heuristic.
With 5 options, 0.50 is a strong plurality; with 2 options, 0.50 is a coin flip.
**Fix:** gate a `Choice` on margin over the runner-up, which is scale-free:
```python
top = probabilities[choice]
runner = max(v for k, v in probabilities.items() if k != choice)
act = top >= 0.45 and (top - runner) >= 0.20
```
`Noul` has no confidence field, so gate it on distance from 0.5:
```python
act = abs(noul - 0.5) >= 0.15 # act when noul >= 0.65 or <= 0.35
```
## Design rules that follow from the above
1. **Never** route an arithmetic comparison through Jev. Compute it in code.
2. Send **conclusions and buckets**, not raw numbers to be compared.
3. Keep `state` small and relevant — context rot is real and measurable.
4. Batch all questions for one state into **one** call.
5. Questions in one call are independent; if Q2 needs Q1's answer, make a
second call.
6. Gate `Choice` on margin, not confidence.
7. Treat a low-confidence answer as a reason to fall back, not to guess.
## Credential handling
The key lives in a sops-nix managed file, **not** in the shell environment:
```
~/.config/secrets/global-env/TYPESAFEAI_API_KEY
-> ~/.config/sops-nix/secrets/TYPESAFEAI_API_KEY
```
`jev.py` reads it at runtime and never logs it. `JevClient.__repr__` prints
`key=REDACTED`, so the secret cannot leak through a traceback or log line.
---
# Session 2 additions
## Option count dilutes a Choice — measured twice
### Shop, 14 candidates
A shop offered 14 affordable items in one `Choice`. Jev's top pick scored only
**0.26** (runner 0.17, margin 0.09) — the probability mass spread across all
fourteen. The margin gate correctly rejected it, so the bot would always leave
the shop with gold unspent.
**Fix:** use the documented **re-ranking** pattern — one *absolute* `Noul` per
candidate, then take the argmax in code. Absolute judgements do not dilute as
the candidate count grows.
### Same problem, smaller: 5 cards
Already covered above. With 5 options, `confidence` is `(5 × top 1) / 4`, so
a clear plurality reads as 0.50. Gate on margin, not confidence.
## The question framing matters more than the threshold
Asking Jev to weigh value against a **number** degrades its judgement, exactly
as the jaggedness page predicts. Measured on the same shop, same state:
| Framing | Spread across candidates | Top item |
|---|---|---|
| "worth its **72 gold** price for this deck?" | **0.28** | Bag of Preparation 0.51 |
| "would this make this deck stronger?" | **0.48** | Bag of Preparation 0.67 |
| "does this fit what this deck is doing?" | 0.51 | Ashen Strike 0.71 |
| "improve more than it dilutes?" | 0.42 | Bag of Preparation 0.65 |
Putting the price in the question **halved the spread** and pulled the top item
below any usable threshold.
**Rule:** filter affordability in code, keep the price in the `state` for
context, and keep it **out of the question**. Ask about deck fit only.
## Risk is not a preference, and Jev is bad at spotting danger
On an event offering "Keep Deciphering" and "Lose Everything":
| Option | "Does this risk losing the run?" |
|---|---|
| "Lose Everything" | **0.46** |
| "Keep Deciphering" | 0.52 |
| "Stop" | 0.38 |
The model ranked the run-ending option as less risky than a moderate one. Do
not use a model judgement to detect danger. Use deterministic keyword matching
plus a stricter confidence gate. See [05](05-failure-modes.md) §10.
## Confirmed working patterns
| Pattern | Where used | Result |
|---|---|---|
| Batch all questions for one state in one call | combat | 0.73 s for 1 question, 0.90 s for 3 |
| Absolute `Noul` per candidate, argmax in code | shop | Works where a 14-way Choice failed |
| Compute the hard fact in code, let Jev pick | lethal, potions | 33 lethal lines executed without the model |
| Gate on margin, not confidence | every `Choice` | Fixed a rejected-correct-answer bug |
| Keep arithmetic out of the question | shop | Doubled the usable spread |
| Deterministic safety net for danger | events | Caught what the model missed |