AI Scientist Bench
All configurations / model

Sonnet 5

2/9 dimensions tested
Sonnet 5
Configuration
Sonnet 5 with no tools and no skill; the floor.
Source
claude-sonnet-5
Recorded runs
715 scored outputs · $4.12 measured model spend · runs on 2026-08-31, 2026-09-15
Access
closed-API model · no tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
8.47Verify
Baselinen=300 · verdict accuracy · evidence in hand · $0.0036/task · 2026-08-31

Run provenance

Report path: outputs/scifact/results_matrix_sonnet5.json
Task pack: scifact-dev-300
Run: matrix-sonnet-5

  • from memory5.69 · n=299 · $0.0050 · 2026-08-31 ·
  • campaign-v310.00 · n=8 · $0.0077 · 2026-09-15 · headline; 8 tasks, 8 source clusters; 0 failed attempts kept in the denominator
Ground

Not yet tested.

Find

Not yet tested.

3.85Extract
Baselinen=100 · mean(entity F1, relation F1) · $0.0138/task · 2026-08-31

entity F1 0.530, relation F1 0.240

Run provenance

Report path: outputs/extract/results.json
Task pack: scierc-test-100
Run: extract-100

  • campaign-v32.40 · n=8 · $0.0124 · 2026-09-15 · headline; 8 tasks, 8 source clusters; 0 failed attempts kept in the denominator
Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.