AI Scientist Bench

All configurations

About the scores ↗

An archive of configurations tested on different task sets. For a skill’s effect over its matching baseline, use the matched results.

Benchmark v0.1·39 configurations with runs·5/9 dimensions scored·Last test 2026-09-15
Tested Serving failure Pending
Benchmark results

Scores /10 · Each column has its own metric · — Untested · Select a score for run details

ConfigurationCost / taskTime / taskTestedVerifyGroundFindExtractAnalyzeComputeSynthesizeReviewHypothesize
Tested At least one scored dimension
short checklistSonnet 5 · Skill
$0.006–
$0.031
3–13s4/910.004.383.0510.00
$0.002–
$0.002
2/98.432.62
$0.001–
$0.009
2/98.409.17canon0.71tail
Opus 5Model
$0.017–
$0.030
2/98.304.48
$0.004–
$0.014
2/98.473.85
BM25 + rerankOpus 4.8 · Harness
$0.00531/98.69
BM25 top-20 + rerankSonnet 5 · Harness
$0.03023.1s1/94.38
BM25 top-5Harness
$0.00001/98.03
calibrated-abstentionHaiku 4.5 · Skill
$0.00251/98.23
calibrated-abstentionSonnet 4.6 · Skill
$0.00611/98.07
chain-of-verificationHaiku 4.5 · Skill
$0.00361/98.20
chain-of-verificationSonnet 4.6 · Skill
$0.00751/98.37
cited abstract suppliedHaiku 4.5 · Harness
$0.00151/98.43
cited abstract suppliedOpus 4.8 · Harness
$0.00901/98.40
cited abstract suppliedOpus 5 · Harness
$0.01661/98.30
cited abstract suppliedSonnet 5 · Harness
$0.00361/98.47
claim-decompositionHaiku 4.5 · Skill
$0.00251/98.23
claim-decompositionSonnet 4.6 · Skill
$0.00691/97.73
Crossref findOpus 4.8 · Harness
$0.005–
$0.005
1/95.83canon10.00tail
expert-personaHaiku 4.5 · Skill
$0.00241/98.30
expert-personaSonnet 4.6 · Skill
$0.00411/98.27
Fable 5Model
$0.017–
$0.060
1/9Failed4.69
K-Dense literature-reviewSonnet 5 · Skill
$0.03363.6s1/94.38
$0.00763.2s1/910.00
$0.027717.4s1/910.00
local SciERC procedureSonnet 5 · Skill
$0.016510.3s1/93.34
pdf-extractionHaiku 4.5 · Skill
$0.00231/98.33
pdf-extractionSonnet 4.6 · Skill
$0.00541/98.17
sandboxed analysisSonnet 5 · Harness
$0.029020.3s1/98.75
schema promptFable 5 · Skill
$0.10151/94.90
schema promptHaiku 4.5 · Skill
$0.00251/93.03
schema promptOpus 5 · Skill
$0.03861/94.74
schema promptSonnet 5 · Skill
$0.01571/93.92
$0.00351/98.00
step-backHaiku 4.5 · Skill
$0.00241/98.20
step-backSonnet 4.6 · Skill
$0.00731/98.00
systematic-reviewHaiku 4.5 · Skill
$0.00271/98.27
systematic-reviewSonnet 4.6 · Skill
$0.00781/98.40
Serving failure A recorded attempt with no score
cited abstract suppliedFable 5 · Harness
$0.01670/9Failed
Pending No benchmark runs yet
ElicitProduct · Pending
0/9
ConsensusProduct · Pending
0/9
SciteProduct · Pending
0/9
UndermindProduct · Pending
0/9
FutureHouse PlatformProduct · Pending
0/9
Perplexity AcademicProduct · Pending
0/9
Edison KosmosProduct · Pending
0/9
SciSpaceProduct · Pending
0/9
OpenAI deep researchProduct · Pending
0/9
Gemini deep researchProduct · Pending
0/9
Claude researchProduct · Pending
0/9
Semantic ScholarProduct · Pending
0/9
PaperQA2Framework · Pending
0/9
STORMFramework · Pending
0/9
gpt-researcherFramework · Pending
0/9
Sakana AI Scientist v2Framework · Pending
0/9
RD-AgentFramework · Pending
0/9
Agent LaboratoryFramework · Pending
0/9
BiomniFramework · Pending
0/9
karpathy/autoresearchFramework · Pending
0/9

Each row is a specific model and configuration. Task sets and sample sizes vary; open a score before comparing results. Ground includes canonical and long-tail claims separately. Cost and time show ranges across the displayed runs.

How scoring works