All configurations
About the scores ↗An archive of configurations tested on different task sets. For a skill’s effect over its matching baseline, use the matched results.
Benchmark v0.1·39 configurations with runs·5/9 dimensions scored·Last test 2026-09-15
Tested Serving failure Pending
| Configuration | Cost / task | Time / task | Tested | Verify | Ground | Find | Extract | Analyze | Compute | Synthesize | Review | Hypothesize |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Tested At least one scored dimension | ||||||||||||
short checklistSonnet 5 · Skill | $0.006– $0.031 | 3–13s | 4/9 | 10.00 | — | 4.38 | 3.05 | 10.00 | — | — | — | — |
Haiku 4.5Model | $0.002– $0.002 | — | 2/9 | 8.43 | — | — | 2.62 | — | — | — | — | — |
Opus 4.8Model | $0.001– $0.009 | — | 2/9 | 8.40 | 9.17canon0.71tail | — | — | — | — | — | — | — |
Opus 5Model | $0.017– $0.030 | — | 2/9 | 8.30 | — | — | 4.48 | — | — | — | — | — |
Sonnet 5Model | $0.004– $0.014 | — | 2/9 | 8.47 | — | — | 3.85 | — | — | — | — | — |
BM25 + rerankOpus 4.8 · Harness | $0.0053 | — | 1/9 | — | — | 8.69 | — | — | — | — | — | — |
BM25 top-20 + rerankSonnet 5 · Harness | $0.0302 | 3.1s | 1/9 | — | — | 4.38 | — | — | — | — | — | — |
BM25 top-5Harness | $0.0000 | — | 1/9 | — | — | 8.03 | — | — | — | — | — | — |
calibrated-abstentionHaiku 4.5 · Skill | $0.0025 | — | 1/9 | 8.23 | — | — | — | — | — | — | — | — |
calibrated-abstentionSonnet 4.6 · Skill | $0.0061 | — | 1/9 | 8.07 | — | — | — | — | — | — | — | — |
chain-of-verificationHaiku 4.5 · Skill | $0.0036 | — | 1/9 | 8.20 | — | — | — | — | — | — | — | — |
chain-of-verificationSonnet 4.6 · Skill | $0.0075 | — | 1/9 | 8.37 | — | — | — | — | — | — | — | — |
cited abstract suppliedHaiku 4.5 · Harness | $0.0015 | — | 1/9 | 8.43 | — | — | — | — | — | — | — | — |
cited abstract suppliedOpus 4.8 · Harness | $0.0090 | — | 1/9 | 8.40 | — | — | — | — | — | — | — | — |
cited abstract suppliedOpus 5 · Harness | $0.0166 | — | 1/9 | 8.30 | — | — | — | — | — | — | — | — |
cited abstract suppliedSonnet 5 · Harness | $0.0036 | — | 1/9 | 8.47 | — | — | — | — | — | — | — | — |
claim-decompositionHaiku 4.5 · Skill | $0.0025 | — | 1/9 | 8.23 | — | — | — | — | — | — | — | — |
claim-decompositionSonnet 4.6 · Skill | $0.0069 | — | 1/9 | 7.73 | — | — | — | — | — | — | — | — |
Crossref findOpus 4.8 · Harness | $0.005– $0.005 | — | 1/9 | — | 5.83canon10.00tail | — | — | — | — | — | — | — |
expert-personaHaiku 4.5 · Skill | $0.0024 | — | 1/9 | 8.30 | — | — | — | — | — | — | — | — |
expert-personaSonnet 4.6 · Skill | $0.0041 | — | 1/9 | 8.27 | — | — | — | — | — | — | — | — |
Fable 5Model | $0.017– $0.060 | — | 1/9 | Failed | — | — | 4.69 | — | — | — | — | — |
K-Dense literature-reviewSonnet 5 · Skill | $0.0336 | 3.6s | 1/9 | — | — | 4.38 | — | — | — | — | — | — |
K-Dense scientific-critical-thinkingSonnet 5 · Skill | $0.0076 | 3.2s | 1/9 | 10.00 | — | — | — | — | — | — | — | — |
K-Dense statistical-analysisSonnet 5 · Skill | $0.0277 | 17.4s | 1/9 | — | — | — | — | 10.00 | — | — | — | — |
local SciERC procedureSonnet 5 · Skill | $0.0165 | 10.3s | 1/9 | — | — | — | 3.34 | — | — | — | — | — |
pdf-extractionHaiku 4.5 · Skill | $0.0023 | — | 1/9 | 8.33 | — | — | — | — | — | — | — | — |
pdf-extractionSonnet 4.6 · Skill | $0.0054 | — | 1/9 | 8.17 | — | — | — | — | — | — | — | — |
sandboxed analysisSonnet 5 · Harness | $0.0290 | 20.3s | 1/9 | — | — | — | — | 8.75 | — | — | — | — |
schema promptFable 5 · Skill | $0.1015 | — | 1/9 | — | — | — | 4.90 | — | — | — | — | — |
schema promptHaiku 4.5 · Skill | $0.0025 | — | 1/9 | — | — | — | 3.03 | — | — | — | — | — |
schema promptOpus 5 · Skill | $0.0386 | — | 1/9 | — | — | — | 4.74 | — | — | — | — | — |
schema promptSonnet 5 · Skill | $0.0157 | — | 1/9 | — | — | — | 3.92 | — | — | — | — | — |
Sonnet 4.6Model | $0.0035 | — | 1/9 | 8.00 | — | — | — | — | — | — | — | — |
step-backHaiku 4.5 · Skill | $0.0024 | — | 1/9 | 8.20 | — | — | — | — | — | — | — | — |
step-backSonnet 4.6 · Skill | $0.0073 | — | 1/9 | 8.00 | — | — | — | — | — | — | — | — |
systematic-reviewHaiku 4.5 · Skill | $0.0027 | — | 1/9 | 8.27 | — | — | — | — | — | — | — | — |
systematic-reviewSonnet 4.6 · Skill | $0.0078 | — | 1/9 | 8.40 | — | — | — | — | — | — | — | — |
| Serving failure A recorded attempt with no score | ||||||||||||
cited abstract suppliedFable 5 · Harness | $0.0167 | — | 0/9 | Failed | — | — | — | — | — | — | — | — |
| Pending No benchmark runs yet | ||||||||||||
ElicitProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
ConsensusProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
SciteProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
UndermindProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
FutureHouse PlatformProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
Perplexity AcademicProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
Edison KosmosProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
SciSpaceProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
OpenAI deep researchProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
Gemini deep researchProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
Claude researchProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
Semantic ScholarProduct · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
PaperQA2Framework · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
STORMFramework · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
gpt-researcherFramework · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
Sakana AI Scientist v2Framework · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
RD-AgentFramework · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
Agent LaboratoryFramework · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
BiomniFramework · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
karpathy/autoresearchFramework · Pending | — | — | 0/9 | — | — | — | — | — | — | — | — | — |
Each row is a specific model and configuration. Task sets and sample sizes vary; open a score before comparing results. Ground includes canonical and long-tail claims separately. Cost and time show ranges across the displayed runs.
How scoring works
- Scores use the column's own metric. An 8.4 in Verify is 84% verdict accuracy. A 3.3 in Extract is F1 0.33. We do not average across dimensions.
- Run conditions matter. Verify displays evidence-in-hand results. Ground keeps canonical and long-tail splits separate. Profiles retain the other conditions and earlier runs.
- Untested means no result. A dash is never a zero. A serving failure is shown separately.
- These are development results. Public task sets can be contaminated, sample sizes vary, and exploratory p-values are not corrected for multiple comparisons. They do not establish a general winner.
- Displayed runs are selected by sample size, then recency. Use head-to-head comparisons for matching task sets. Scores and sample sizes remain available in the downloadable results.
- Costs come from the recorded runs. Missing timings stay blank. Raw report paths are included as provenance; the full reports are not bundled with this site.