AI Scientist Bench
All configurations / harness

cited abstract supplied

1/9 dimensions tested
Opus 5
Configuration
The claim's cited abstract is handed to the model before it answers. No search, no tools; the evidence document itself.
Source
gym.py scifact · contestants V_ev_*
Recorded runs
283 scored outputs · $4.69 measured model spend · runs on 2026-08-31
Access
open harness · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
8.30Verify
Compared with baselinen=283 · verdict accuracy · evidence in hand · Δ +2.60 vs Opus 5 · 85W 13L 179T on 277 paired · p<.001 (unadjusted) · $0.0166/task · 2026-08-31

28 of 600 calls lost to 529 Overloaded; dropped, not scored; naked 0.578 at $0.0278 per task; the harness is cheaper because the naked model burns tokens hedging Test: exact McNemar on the tasks both cells answered.

Run provenance

Report path: outputs/scifact/results_matrix_opus5.json
Task pack: scifact-dev-300
Run: matrix-opus-5

Ground

Not yet tested.

Find

Not yet tested.

Extract

Not yet tested.

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.