AI Scientist Bench
All configurations / prompt skill

schema prompt

1/9 dimensions tested
Haiku 4.5
Configuration
The SciERC type definitions and output schema written into the prompt; zero-shot otherwise.
Source
contestants.py E1_*
Recorded runs
100 scored outputs · $0.25 measured model spend · runs on 2026-08-31
Access
open prompt · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
Verify

Not yet tested.

Ground

Not yet tested.

Find

Not yet tested.

3.03Extract
Compared with baselinen=100 · mean(entity F1, relation F1) · Δ +0.41 vs Haiku 4.5 · 68W 30L 2T on 100 paired · p<.001 (unadjusted) · $0.0025/task · 2026-08-31

entity F1 0.476, relation F1 0.130 Test: exact sign test on discordant pairs.

Run provenance

Report path: outputs/extract/results.json
Task pack: scierc-test-100
Run: extract-100

Analyze

Not yet tested.

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.