AI Scientist Bench
All configurations / prompt skill

short checklist

4/9 dimensions tested
Sonnet 5
Configuration
A short, track-specific checklist prepended to the baseline prompt.
Source
campaign/run.py
Recorded runs
31 scored outputs · $0.49 measured model spend · runs on 2026-09-15
Access
open prompt · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
10.00Verify
Compared with baselinen=8 · verdict accuracy · Δ ±0.00 vs Sonnet 5 · 0W 0L 8T on 8 paired · p=1.000 (unadjusted) · $0.0059/task · 3.1s/task · 2026-09-15

headline; 8 tasks, 8 source clusters; 0 failed attempts kept in the denominator Test: exact sign test on discordant pairs; n too small for a winner.

Run provenance

Report path: outputs/campaign/20260915/summary.json
Task pack: campaign-dev-20260915
Run: campaign-v3

Ground

Not yet tested.

4.38Find
Compared with baselinen=4 · recall@5 · Δ ±0.00 vs BM25 top-20 + rerank · 0W 0L 4T on 4 paired · p=1.000 (unadjusted) · $0.0314/task · 4.0s/task · 2026-09-15

headline; 4 tasks, 4 source clusters; 0 failed attempts kept in the denominator; free BM25 0.354; candidate oracle 0.500; 2 of 4 pools reachable Test: exact sign test on discordant pairs; n too small for a winner.

Run provenance

Report path: outputs/campaign/20260915/summary.json
Task pack: campaign-dev-20260915
Run: campaign-v3

  • targeted recoverable misses10.00 · n=2 · Δ ±0.00 · p=1.000 (unadjusted) · $0.0250 · 2026-09-15 · targeted recoverable misses; 2 tasks, 2 source clusters; 0 failed attempts kept in the denominator; free BM25 0.000; candidate oracle 1.000; 2 of 2 pools reachable
  • retrieval-ceiling control10.00 · n=1 · Δ ±0.00 · p=1.000 (unadjusted) · $0.0261 · 2026-09-15 · retrieval-ceiling control; 1 tasks, 1 source clusters; 0 failed attempts kept in the denominator; free BM25 1.000; candidate oracle 1.000; 1 of 1 pools reachable
3.05Extract
Compared with baselinen=8 · mean(entity F1, relation F1) · Δ +0.66 vs Sonnet 5 · 5W 1L 2T on 8 paired · p=.219 (unadjusted) · $0.0131/task · 7.5s/task · 2026-09-15

headline; 8 tasks, 8 source clusters; 1 failed attempts kept in the denominator Test: exact sign test on discordant pairs; n too small for a winner.

Run provenance

Report path: outputs/campaign/20260915/summary.json
Task pack: campaign-dev-20260915
Run: campaign-v3

10.00Analyze
Compared with baselinen=8 · executed pass rate, original and perturbed inputs · Δ +1.25 vs sandboxed analysis · 1W 0L 7T on 8 paired · p=1.000 (unadjusted) · $0.0166/task · 13.1s/task · 2026-09-15

headline; 8 tasks, 5 source clusters; 0 failed attempts kept in the denominator Test: exact sign test on discordant pairs; n too small for a winner.

Run provenance

Report path: outputs/campaign/20260915/summary.json
Task pack: campaign-dev-20260915
Run: campaign-v3

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.