All configurations / harness
sandboxed analysis
- Configuration
- The model writes and runs code in a Seatbelt sandbox (NumPy, SciPy, pandas) with the baseline prompt.
- Source
- campaign/analysis_sandbox.py
- Recorded runs
- 8 scored outputs · $0.23 measured model spend · runs on 2026-09-15
- Access
- open harness · standard tools
FIG. 01
By dimension
Scores /10 · Paired p-values are exploratory and unadjusted—Verify
Not yet tested.
—Ground
Not yet tested.
—Find
Not yet tested.
—Extract
Not yet tested.
8.75Analyze
Baselinen=8 · executed pass rate, original and perturbed inputs · $0.0290/task · 20.3s/task · 2026-09-15
headline; 8 tasks, 5 source clusters; 1 failed attempts kept in the denominator; the failure is a tool call that exhausted its output tokens without a code argument, a workflow failure not a wrong regression
Run provenance
Report path: outputs/campaign/20260915/summary.json
Task pack: campaign-dev-20260915
Run: campaign-v3
—Compute
Not yet tested.
—Synthesize
Not yet tested.
—Review
Not yet tested.
—Hypothesize
Not yet tested.