AI Scientist Bench
All configurations / harness

Crossref find

1/9 dimensions tested
Opus 4.8
Configuration
A keyword bibliographic search on Crossref; the model picks one DOI from the real hits, so it cannot fabricate one. Does not index arXiv.
Source
contestants.py L1_crossref
Recorded runs
26 scored outputs · $0.12 measured model spend · runs on 2026-07-13
Access
open harness · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
Verify

Not yet tested.

5.83Ground · canonical
Compared with baselinen=12 · right DOI · split canonical · Δ −3.34 vs Opus 4.8 · 0W 4L 8T on 12 paired · p=.125 (unadjusted) · $0.0048/task · 2026-07-13

12 memorized-canon claims; naked BEATS harness here; grounding 0.83, hallucination 0.17 Test: exact McNemar on discordant pairs.

Run provenance

Report path: outputs/grounding/results.json
Task pack: grounding-canonical
Run: grounding

10.00Ground · tail
Compared with baselinen=14 · right DOI · split tail · Δ +9.29 vs Opus 4.8 · 13W 0L 1T on 14 paired · p<.001 (unadjusted) · $0.0046/task · 2026-07-13

14 niche 2024-25 claims; harness EARNS its keep here; grounding 1.00, hallucination 0.00 Test: exact McNemar on discordant pairs.

Run provenance

Report path: outputs/grounding-recent/results.json
Task pack: grounding-tail
Run: grounding-recent

Find

Not yet tested.

Extract

Not yet tested.

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.