AI Scientist Bench
All configurations / model

Fable 5

1/9 dimensions tested
Fable 5
Configuration
Fable 5 with no tools and no skill; the floor.
Source
claude-fable-5
Recorded runs
700 scored outputs · $16.09 measured model spend · runs on 2026-08-31
Access
closed-API model · no tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
voidVerify
Serving failuren=300 · verdict accuracy · evidence in hand · $0.0167/task · 2026-08-31

two of every three calls returned zero output tokens; a serving failure, not a score

Run provenance

Report path: outputs/scifact/results_matrix_fable5.json
Task pack: scifact-dev-300
Run: matrix-fable-5

  • from memoryvoid · n=300 · $0.0169 · 2026-08-31 · two of every three calls returned zero output tokens; a serving failure, not a score
Ground

Not yet tested.

Find

Not yet tested.

4.69Extract
Baselinen=100 · mean(entity F1, relation F1) · $0.0602/task · 2026-08-31

entity F1 0.586, relation F1 0.352

Run provenance

Report path: outputs/extract/results.json
Task pack: scierc-test-100
Run: extract-100

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.