AI Scientist Bench
All configurations / model

Haiku 4.5

2/9 dimensions tested
Haiku 4.5
Configuration
Haiku 4.5 with no tools and no skill; the floor.
Source
claude-haiku-4-5
Recorded runs
1,300 scored outputs · $2.41 measured model spend · runs on 2026-07-14, 2026-08-31
Access
closed-API model · no tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
8.43Verify
Baselinen=300 · verdict accuracy · evidence in hand · $0.0015/task · 2026-08-31

Run provenance

Report path: outputs/scifact/results_matrix_haiku45.json
Task pack: scifact-dev-300
Run: matrix-haiku-4.5

  • from memory4.60 · n=300 · $0.0019 · 2026-08-31 ·
  • evidence in hand8.20 · n=300 · $0.0019 · 2026-07-14 · the arena floor; NEI class 0.72
  • from memory4.87 · n=300 · $0.0019 · 2026-07-14 · the arena floor; NEI class 0.20
Ground

Not yet tested.

Find

Not yet tested.

2.62Extract
Baselinen=100 · mean(entity F1, relation F1) · $0.0023/task · 2026-08-31

entity F1 0.404, relation F1 0.120

Run provenance

Report path: outputs/extract/results.json
Task pack: scierc-test-100
Run: extract-100

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.