All configurations / prompt skill
chain-of-verification
- Configuration
- Draft a verdict, generate verification questions, answer them against the evidence, revise. Dhuliawala et al. 2023.
- Source
- gym/skills/chain-of-verification.md · public · Dhuliawala et al. 2023
- Recorded runs
- 600 scored outputs · $2.05 measured model spend · runs on 2026-07-14
- Access
- open prompt · standard tools
FIG. 01
By dimension
Scores /10 · Paired p-values are exploratory and unadjusted8.20Verify
Compared with baselinen=300 · verdict accuracy · evidence in hand · Δ ±0.00 vs Haiku 4.5 · 14W 14L 272T on 300 paired · p=1.000 (unadjusted) · $0.0036/task · 2026-07-14
NEI class 0.77 vs naked 0.72; 575 output tokens per call Test: exact McNemar on discordant pairs.
Run provenance
Report path: outputs/arena/scifact_haiku_evidence.json
Task pack: scifact-dev-300
Run: arena-haiku-4.5
- from memory4.17 · n=300 · Δ −0.70 · p=.005 (unadjusted) · $0.0033 · 2026-07-14 · NEI class 0.19 vs naked 0.20; 613 output tokens per call
—Ground
Not yet tested.
—Find
Not yet tested.
—Extract
Not yet tested.
—Analyze
Not yet tested.
—Compute
Not yet tested.
—Synthesize
Not yet tested.
—Review
Not yet tested.
—Hypothesize
Not yet tested.