All configurations / prompt skill
chain-of-verification
- Configuration
- Draft a verdict, generate verification questions, answer them against the evidence, revise. Dhuliawala et al. 2023.
- Source
- gym/skills/chain-of-verification.md · public · Dhuliawala et al. 2023
- Recorded runs
- 600 scored outputs · $5.01 measured model spend · runs on 2026-07-14
- Access
- open prompt · standard tools
FIG. 01
By dimension
Scores /10 · Paired p-values are exploratory and unadjusted8.37Verify
Compared with baselinen=300 · verdict accuracy · evidence in hand · Δ +0.37 vs Sonnet 4.6 · 15W 4L 281T on 300 paired · p=.019 (unadjusted) · $0.0075/task · 2026-07-14
NEI class 0.77 vs naked 0.70; 367 output tokens per call Test: exact McNemar on discordant pairs.
Run provenance
Report path: outputs/arena/scifact_sonnet_evidence.json
Task pack: scifact-dev-300
Run: arena-sonnet-4.6
- from memory4.67 · n=300 · Δ −0.30 · p=.222 (unadjusted) · $0.0091 · 2026-07-14 · NEI class 0.05 vs naked 0.10; 564 output tokens per call
—Ground
Not yet tested.
—Find
Not yet tested.
—Extract
Not yet tested.
—Analyze
Not yet tested.
—Compute
Not yet tested.
—Synthesize
Not yet tested.
—Review
Not yet tested.
—Hypothesize
Not yet tested.