All configurations / prompt skill
calibrated-abstention
- Configuration
- Prefer NEI unless the evidence directly addresses the claim.
- Source
- gym/skills/calibrated-abstention.md · public · research-prompt
- Recorded runs
- 600 scored outputs · $3.79 measured model spend · runs on 2026-07-14
- Access
- open prompt · standard tools
FIG. 01
By dimension
Scores /10 · Paired p-values are exploratory and unadjusted8.07Verify
Compared with baselinen=300 · verdict accuracy · evidence in hand · Δ +0.07 vs Sonnet 4.6 · 6W 4L 290T on 300 paired · p=.754 (unadjusted) · $0.0061/task · 2026-07-14
NEI class 0.71 vs naked 0.70; 272 output tokens per call Test: exact McNemar on discordant pairs.
Run provenance
Report path: outputs/arena/scifact_sonnet_evidence.json
Task pack: scifact-dev-300
Run: arena-sonnet-4.6
- from memory4.90 · n=300 · Δ −0.07 · p=.856 (unadjusted) · $0.0065 · 2026-07-14 · NEI class 0.18 vs naked 0.10; 393 output tokens per call
—Ground
Not yet tested.
—Find
Not yet tested.
—Extract
Not yet tested.
—Analyze
Not yet tested.
—Compute
Not yet tested.
—Synthesize
Not yet tested.
—Review
Not yet tested.
—Hypothesize
Not yet tested.