# 02 — System One / Jev Everything here is measured against `jev-latest` → `jev-1.13.0` from `https://api.typesafe.ai/v1/systemone`. ## What Jev is TypeSafe states it plainly: > "System One is TypeSafe's model for building AI-powered software, not agents. > It does not generate code or choose its own next action." It is a **calibrated classifier**. It evaluates a `state` and answers typed questions about it. It cannot plan, cannot call tools, and cannot invent values outside the option set you supply. | Primitive | Shape | Returns | |---|---|---| | `Choice` | pick 1 of ≤ 255 options | `choice`, `probabilities`, `confidence` | | `Score` | position on ordered levels | `score`, `legend`, `probabilities`, `confidence` | | `Noul` | yes/no | `noul` (probability only, **no confidence**) | Documented properties: questions in one call are evaluated **independently and in parallel**. One answer is never hidden context for another. ## Measured latency — 7× the documented figure The docs say "most queries complete in about 100 ms". Measured end-to-end from this machine: | Request | Time | |---|---| | Minimal (1 short noul) | 0.73 s | | Minimal × 5 repeats | 0.73, 0.73, 0.74, 0.80, 0.76 s | | Full combat state, 3 questions | 0.90 s | | `GET /v1/models` TTFB | 0.68 s (connect 0.22 s) | The floor is **~0.73 s**, dominated by network RTT plus server TTFB. **The useful consequence:** a 1-question call and a 3-question call over a full combat state differ by only 0.17 s. State size and question count are nearly free. **Batch every question for a state into one call.** ## Cost | | | |---|---| | Input | $0.042 per Mtok | | Output | free | | Rate limit | 250k tok/s, 1,200 req/min | | Context | 64k total; 32k for `state` + longest question | | Input type | text only (string, JSON object, array) | For comparison, the STS2MCP README reports a full run costs a frontier LLM about **8M tokens** (~$20–40). The same volume through Jev is roughly **$0.08–0.35 per run**. ## THE CRITICAL FINDING: Jev cannot do arithmetic Ground truth: `energy 3`, hand `Strike(1 cost, 6 dmg) ×3` plus `Bash(2 cost, 8 dmg)`, target on `19 HP`. Maximum reachable damage is **18**, so lethal is **NO**. | Question | Jev answered | Verdict | |---|---|---| | Max damage bucket | `18_to_23` @ 0.73 | **Correct** | | Is lethal available? | `0.79` | **Wrong** | Jev bucketed the magnitude correctly but failed the threshold comparison of 18 versus 19 — **and reported 0.79 confidence on the wrong answer**. Two conclusions: 1. Jev is decent at *approximate magnitude*, poor at *exact comparison*. 2. **Confidence gating cannot protect against arithmetic errors.** A 0.79-confidence wrong answer passes any sane threshold. This is why `facts.py` computes every sum, comparison, and threshold, and hands Jev only conclusions (`"lethal_available": true`). This is a regression test: see `test_facts.py` case 1. The jaggedness page confirms the general shape: Jev "does not count reliably", "is not a calculator", degrades with indirection, and suffers "context rot" when the state carries irrelevant detail. ## The confidence gate trap **A fixed confidence floor is wrong when the option count varies.** Measured: 5 cards offered. Jev picks `Bash` at **0.61**, with `Defend` at 0.29. Reported `confidence` is **0.50**. Why: confidence measures *peakedness*. For 5 options the documented formula is `(count × peak − 1) / (count − 1)` = `(5 × 0.61 − 1) / 4` = 0.50. A 0.55 floor rejected a clear plurality and fell back to a worse heuristic. With 5 options, 0.50 is a strong plurality; with 2 options, 0.50 is a coin flip. **Fix:** gate a `Choice` on margin over the runner-up, which is scale-free: ```python top = probabilities[choice] runner = max(v for k, v in probabilities.items() if k != choice) act = top >= 0.45 and (top - runner) >= 0.20 ``` `Noul` has no confidence field, so gate it on distance from 0.5: ```python act = abs(noul - 0.5) >= 0.15 # act when noul >= 0.65 or <= 0.35 ``` ## Design rules that follow from the above 1. **Never** route an arithmetic comparison through Jev. Compute it in code. 2. Send **conclusions and buckets**, not raw numbers to be compared. 3. Keep `state` small and relevant — context rot is real and measurable. 4. Batch all questions for one state into **one** call. 5. Questions in one call are independent; if Q2 needs Q1's answer, make a second call. 6. Gate `Choice` on margin, not confidence. 7. Treat a low-confidence answer as a reason to fall back, not to guess. ## Credential handling The key lives in a sops-nix managed file, **not** in the shell environment: ``` ~/.config/secrets/global-env/TYPESAFEAI_API_KEY -> ~/.config/sops-nix/secrets/TYPESAFEAI_API_KEY ``` `jev.py` reads it at runtime and never logs it. `JevClient.__repr__` prints `key=REDACTED`, so the secret cannot leak through a traceback or log line. --- # Session 2 additions ## Option count dilutes a Choice — measured twice ### Shop, 14 candidates A shop offered 14 affordable items in one `Choice`. Jev's top pick scored only **0.26** (runner 0.17, margin 0.09) — the probability mass spread across all fourteen. The margin gate correctly rejected it, so the bot would always leave the shop with gold unspent. **Fix:** use the documented **re-ranking** pattern — one *absolute* `Noul` per candidate, then take the argmax in code. Absolute judgements do not dilute as the candidate count grows. ### Same problem, smaller: 5 cards Already covered above. With 5 options, `confidence` is `(5 × top − 1) / 4`, so a clear plurality reads as 0.50. Gate on margin, not confidence. ## The question framing matters more than the threshold Asking Jev to weigh value against a **number** degrades its judgement, exactly as the jaggedness page predicts. Measured on the same shop, same state: | Framing | Spread across candidates | Top item | |---|---|---| | "worth its **72 gold** price for this deck?" | **0.28** | Bag of Preparation 0.51 | | "would this make this deck stronger?" | **0.48** | Bag of Preparation 0.67 | | "does this fit what this deck is doing?" | 0.51 | Ashen Strike 0.71 | | "improve more than it dilutes?" | 0.42 | Bag of Preparation 0.65 | Putting the price in the question **halved the spread** and pulled the top item below any usable threshold. **Rule:** filter affordability in code, keep the price in the `state` for context, and keep it **out of the question**. Ask about deck fit only. ## Risk is not a preference, and Jev is bad at spotting danger On an event offering "Keep Deciphering" and "Lose Everything": | Option | "Does this risk losing the run?" | |---|---| | "Lose Everything" | **0.46** | | "Keep Deciphering" | 0.52 | | "Stop" | 0.38 | The model ranked the run-ending option as less risky than a moderate one. Do not use a model judgement to detect danger. Use deterministic keyword matching plus a stricter confidence gate. See [05](05-failure-modes.md) §10. ## Confirmed working patterns | Pattern | Where used | Result | |---|---|---| | Batch all questions for one state in one call | combat | 0.73 s for 1 question, 0.90 s for 3 | | Absolute `Noul` per candidate, argmax in code | shop | Works where a 14-way Choice failed | | Compute the hard fact in code, let Jev pick | lethal, potions | 33 lethal lines executed without the model | | Gate on margin, not confidence | every `Choice` | Fixed a rejected-correct-answer bug | | Keep arithmetic out of the question | shop | Doubled the usable spread | | Deterministic safety net for danger | events | Caught what the model missed |