AI Scientist Bench
All configurations / harness

BM25 top-20 + rerank

1/9 dimensions tested
Sonnet 5
Configuration
A frozen BM25 top-20 pool reranked by the model to exactly five IDs with the baseline prompt.
Source
campaign/verify_find.py
Recorded runs
7 scored outputs · $0.20 measured model spend · runs on 2026-09-15
Access
open harness · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
Verify

Not yet tested.

Ground

Not yet tested.

4.38Find
Baselinen=4 · recall@5 · $0.0302/task · 3.1s/task · 2026-09-15

headline; 4 tasks, 4 source clusters; 0 failed attempts kept in the denominator; free BM25 0.354; candidate oracle 0.500; 2 of 4 pools reachable

Run provenance

Report path: outputs/campaign/20260915/summary.json
Task pack: campaign-dev-20260915
Run: campaign-v3

  • targeted recoverable misses10.00 · n=2 · $0.0251 · 2026-09-15 · targeted recoverable misses; 2 tasks, 2 source clusters; 0 failed attempts kept in the denominator; free BM25 0.000; candidate oracle 1.000; 2 of 2 pools reachable
  • retrieval-ceiling control10.00 · n=1 · $0.0294 · 2026-09-15 · retrieval-ceiling control; 1 tasks, 1 source clusters; 0 failed attempts kept in the denominator; free BM25 1.000; candidate oracle 1.000; 1 of 1 pools reachable
Extract

Not yet tested.

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.