AI Scientist Bench
All configurations / harness

sandboxed analysis

1/9 dimensions tested
Sonnet 5
Configuration
The model writes and runs code in a Seatbelt sandbox (NumPy, SciPy, pandas) with the baseline prompt.
Source
campaign/analysis_sandbox.py
Recorded runs
8 scored outputs · $0.23 measured model spend · runs on 2026-09-15
Access
open harness · standard tools
FIG. 01

By dimension

Scores /10 · Paired p-values are exploratory and unadjusted
Verify

Not yet tested.

Ground

Not yet tested.

Find

Not yet tested.

Extract

Not yet tested.

8.75Analyze
Baselinen=8 · executed pass rate, original and perturbed inputs · $0.0290/task · 20.3s/task · 2026-09-15

headline; 8 tasks, 5 source clusters; 1 failed attempts kept in the denominator; the failure is a tool call that exhausted its output tokens without a code argument, a workflow failure not a wrong regression

Run provenance

Report path: outputs/campaign/20260915/summary.json
Task pack: campaign-dev-20260915
Run: campaign-v3

Compute

Not yet tested.

Synthesize

Not yet tested.

Review

Not yet tested.

Hypothesize

Not yet tested.