All configurations / prompt skill
systematic-review
- Configuration
- List the claim's components, check each against the evidence, rule NEI when the evidence does not address a component.
- Source
- gym/skills/systematic-review.md · public · skill-pack
- Recorded runs
- 600 scored outputs · $1.58 measured model spend · runs on 2026-07-14
- Access
- open prompt · standard tools
FIG. 01
By dimension
Scores /10 · Paired p-values are exploratory and unadjusted8.27Verify
Compared with baselinen=300 · verdict accuracy · evidence in hand · Δ +0.07 vs Haiku 4.5 · 14W 12L 274T on 300 paired · p=.845 (unadjusted) · $0.0027/task · 2026-07-14
NEI class 0.80 vs naked 0.72; 408 output tokens per call Test: exact McNemar on discordant pairs.
Run provenance
Report path: outputs/arena/scifact_haiku_evidence.json
Task pack: scifact-dev-300
Run: arena-haiku-4.5
- from memory4.83 · n=300 · Δ −0.04 · p=1.000 (unadjusted) · $0.0025 · 2026-07-14 · NEI class 0.28 vs naked 0.20; 462 output tokens per call
—Ground
Not yet tested.
—Find
Not yet tested.
—Extract
Not yet tested.
—Analyze
Not yet tested.
—Compute
Not yet tested.
—Synthesize
Not yet tested.
—Review
Not yet tested.
—Hypothesize
Not yet tested.