AI Scientist Bench
All configurations / prompt skill

chain-of-verification

1/9 dimensions tested
Haiku 4.5
Configuration
Draft a verdict, generate verification questions, answer them against the evidence, revise. Dhuliawala et al. 2023.
Source
gym/skills/chain-of-verification.md · public · Dhuliawala et al. 2023
Recorded runs
600 scored outputs · $2.05 measured model spend · runs on 2026-07-14
Access
open prompt · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
8.20Verify
Compared with baselinen=300 · verdict accuracy · evidence in hand · Δ ±0.00 vs Haiku 4.5 · 14W 14L 272T on 300 paired · p=1.000 (unadjusted) · $0.0036/task · 2026-07-14

NEI class 0.77 vs naked 0.72; 575 output tokens per call Test: exact McNemar on discordant pairs.

Run provenance

Report path: outputs/arena/scifact_haiku_evidence.json
Task pack: scifact-dev-300
Run: arena-haiku-4.5

  • from memory4.17 · n=300 · Δ −0.70 · p=.005 (unadjusted) · $0.0033 · 2026-07-14 · NEI class 0.19 vs naked 0.20; 613 output tokens per call
Ground

Not yet tested.

Find

Not yet tested.

Extract

Not yet tested.

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.