AI Scientist Bench
All configurations / harness

BM25 + rerank

1/9 dimensions tested
Opus 4.8
Configuration
BM25 candidates reranked by the model to five.
Source
contestants.py L1_rerank
Recorded runs
60 scored outputs · $0.32 measured model spend · runs on 2026-07-13
Access
open harness · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
Verify

Not yet tested.

Ground

Not yet tested.

8.69Find
Compared with baselinen=60 · recall@5 · Δ +0.66 vs BM25 top-5 · 7W 1L 52T on 60 paired · p=.070 (unadjusted) · $0.0053/task · 2026-07-13

precision@5 0.468 vs 0.177 Test: exact sign test on discordant pairs.

Run provenance

Report path: outputs/find/results.json
Task pack: scifact-retrieval-60
Run: find-60

Extract

Not yet tested.

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.