AI Scientist Bench
All configurations / model

Opus 5

2/9 dimensions tested
Opus 5
Configuration
Opus 5 with no tools and no skill; the floor.
Source
claude-opus-5
Recorded runs
672 scored outputs · $15.75 measured model spend · runs on 2026-08-31
Access
closed-API model · no tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
8.30Verify
Baselinen=283 · verdict accuracy · evidence in hand · $0.0166/task · 2026-08-31

28 of 600 calls lost to 529 Overloaded; dropped, not scored

Run provenance

Report path: outputs/scifact/results_matrix_opus5.json
Task pack: scifact-dev-300
Run: matrix-opus-5

  • from memory5.78 · n=289 · $0.0278 · 2026-08-31 · 28 of 600 calls lost to 529 Overloaded; dropped, not scored
Ground

Not yet tested.

Find

Not yet tested.

4.48Extract
Baselinen=100 · mean(entity F1, relation F1) · $0.0303/task · 2026-08-31

entity F1 0.546, relation F1 0.350

Run provenance

Report path: outputs/extract/results.json
Task pack: scierc-test-100
Run: extract-100

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.