AI Scientist Bench
All configurations / model

Opus 4.8

2/9 dimensions tested
Opus 4.8
Configuration
Opus 4.8 with no tools and no skill; the floor.
Source
claude-opus-4-8
Recorded runs
626 scored outputs · $6.51 measured model spend · runs on 2026-07-13, 2026-08-31
Access
closed-API model · no tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
8.40Verify
Baselinen=300 · verdict accuracy · evidence in hand · $0.0090/task · 2026-08-31

ran at a 512-token cap; 179 naked responses hit it; history row, not comparable to the 4096-cap matrix

Run provenance

Report path: outputs/scifact/results_opus48_n300_2026-07-13.json
Task pack: scifact-dev-300
Run: matrix-opus-4.8

  • from memory4.77 · n=300 · $0.0125 · 2026-08-31 · ran at a 512-token cap; 179 naked responses hit it; history row, not comparable to the 4096-cap matrix
9.17Ground · canonical
Baselinen=12 · right DOI · split canonical · $0.0009/task · 2026-07-13

12 memorized-canon claims; naked BEATS harness here; grounding 1.00, hallucination 0.00

Run provenance

Report path: outputs/grounding/results.json
Task pack: grounding-canonical
Run: grounding

0.71Ground · tail
Baselinen=14 · right DOI · split tail · $0.0033/task · 2026-07-13

14 niche 2024-25 claims; harness EARNS its keep here; grounding 0.43, hallucination 0.57

Run provenance

Report path: outputs/grounding-recent/results.json
Task pack: grounding-tail
Run: grounding-recent

Find

Not yet tested.

Extract

Not yet tested.

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.