AI Scientist Bench
All configurations / model

Sonnet 4.6

1/9 dimensions tested
Sonnet 4.6
Configuration
Sonnet 4.6 with no tools and no skill; the floor.
Source
claude-sonnet-4-6
Recorded runs
600 scored outputs · $2.55 measured model spend · runs on 2026-07-14
Access
closed-API model · no tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
8.00Verify
Baselinen=300 · verdict accuracy · evidence in hand · $0.0035/task · 2026-07-14

the arena floor; NEI class 0.70

Run provenance

Report path: outputs/arena/scifact_sonnet_evidence.json
Task pack: scifact-dev-300
Run: arena-sonnet-4.6

  • from memory4.97 · n=300 · $0.0050 · 2026-07-14 · the arena floor; NEI class 0.10
Ground

Not yet tested.

Find

Not yet tested.

Extract

Not yet tested.

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.