All configurations / model
Opus 4.8
- Configuration
- Opus 4.8 with no tools and no skill; the floor.
- Source
- claude-opus-4-8
- Recorded runs
- 626 scored outputs · $6.51 measured model spend · runs on 2026-07-13, 2026-08-31
- Access
- closed-API model · no tools
FIG. 01
By dimension
Scores /10 · Paired p-values are exploratory and unadjusted8.40Verify
Baselinen=300 · verdict accuracy · evidence in hand · $0.0090/task · 2026-08-31
ran at a 512-token cap; 179 naked responses hit it; history row, not comparable to the 4096-cap matrix
Run provenance
Report path: outputs/scifact/results_opus48_n300_2026-07-13.json
Task pack: scifact-dev-300
Run: matrix-opus-4.8
- from memory4.77 · n=300 · $0.0125 · 2026-08-31 · ran at a 512-token cap; 179 naked responses hit it; history row, not comparable to the 4096-cap matrix
9.17Ground · canonical
Baselinen=12 · right DOI · split canonical · $0.0009/task · 2026-07-13
12 memorized-canon claims; naked BEATS harness here; grounding 1.00, hallucination 0.00
Run provenance
Report path: outputs/grounding/results.json
Task pack: grounding-canonical
Run: grounding
0.71Ground · tail
Baselinen=14 · right DOI · split tail · $0.0033/task · 2026-07-13
14 niche 2024-25 claims; harness EARNS its keep here; grounding 0.43, hallucination 0.57
Run provenance
Report path: outputs/grounding-recent/results.json
Task pack: grounding-tail
Run: grounding-recent
—Find
Not yet tested.
—Extract
Not yet tested.
—Analyze
Not yet tested.
—Compute
Not yet tested.
—Synthesize
Not yet tested.
—Review
Not yet tested.
—Hypothesize
Not yet tested.