AI Scientist Bench

Skill comparisons

All runs ↗
Sonnet 5· 28 tasks · 84 attempts · 21 sources · $1.47· 2026-09-15

Pilot scores

Analysis reproduction ↓
TaskAI alone+ Checklist+ Skill
Check a claim ↗n=8 · verdict accuracy8/8$0.0077 · 3.4s8/8$0.0059 · 3.1s8/8$0.0076 · 3.2s
Choose papers ↗n=4 · recall@5 · ceiling 50%43.8%$0.0302 · 3.1s43.8%$0.0314 · 4.0s43.8%$0.0336 · 3.6s
Extract information ↗n=8 · F1 · local procedure0.240$0.0124 · 7.4s0.306$0.0131 · 7.5s0.334$0.0165 · 10.3s
↳ Wrappers removedPost-hoc rescore · same 8 answers0.4000.3410.391
Analyze data ↗n=8 · 1 baseline tool-call failure7/8$0.0290 · 20.3s8/8$0.0166 · 13.1s8/8$0.0277 · 17.4s

Score, then mean $ and seconds per attempt. One attempt per setup per task. Public development cases.

Analysis reproduction

1 case · 3 runs · $0.39

Sonnet 5 · Claude Code 2.1.272 · K-Dense statistical-analysis ↗

SetupCorrectReplay + figureSkill openedCostElapsed
AI alone3/3$0.1162131.3s
With checklist3/3$0.1607258.9s
Skill installed3/3No$0.1163137.1s
Residual variances, all setupsExam1 118.195Exam2 124.754Exam3 87.973

One supplied-code case. Installed skill never opened. Elapsed time includes file searches.

Source & run details

Reproduce a published structural-equation analysis ↗ · 2026-09-15 · skill revision 330c8e7.

One supplied-code teaching case with one attempt per setup. Three values come from one model fit; they are not three independent cases. The checklist includes a task-specific residual-variance reminder. All setups retain the native host’s built-in skills. Measured benchmark model calls only; local compute, runtime setup, storage and assistant development are excluded.

Download native results ↗

Methodology & limitations

Same model, inputs and limits within each comparison. Failures remain in the denominator. Pilot skills are supplied instructions; the native run tests an installed package. Counts are tasks, not necessarily independent datasets. No overall score.

Claims: verdict agreement with dataset labels. Papers: recall@5 from a fixed list. Extraction: entity/relation F1; malformed output scores zero. Analysis: code must pass original and changed-input checks. The wrapper rescore is a post-hoc sensitivity check with unchanged entity/relation content.

Model calls: pilot table $1.4741; full campaign including calibration and diagnostics $2.8401. Matched results ↗ · Format diagnostic ↗

Next experiments