AI Scientist Bench
All configurations / retriever

BM25 top-5

1/9 dimensions tested
no model
Configuration
Lexical BM25 over the 5,183-abstract corpus, top five returned. No model. The Find floor.
Source
contestants.py L0_bm25
Recorded runs
60 scored outputs · $0.00 measured model spend · runs on 2026-07-13
Access
open harness · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
Verify

Not yet tested.

Ground

Not yet tested.

8.03Find
Baselinen=60 · recall@5 · $0.0000/task · 2026-07-13

precision@5 0.177

Run provenance

Report path: outputs/find/results.json
Task pack: scifact-retrieval-60
Run: find-60

Extract

Not yet tested.

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.