All configurations / harness
Crossref find
- Configuration
- A keyword bibliographic search on Crossref; the model picks one DOI from the real hits, so it cannot fabricate one. Does not index arXiv.
- Source
- contestants.py L1_crossref
- Recorded runs
- 26 scored outputs · $0.12 measured model spend · runs on 2026-07-13
- Access
- open harness · standard tools
FIG. 01
By dimension
Scores /10 · Paired p-values are exploratory and unadjusted—Verify
Not yet tested.
5.83Ground · canonical
Compared with baselinen=12 · right DOI · split canonical · Δ −3.34 vs Opus 4.8 · 0W 4L 8T on 12 paired · p=.125 (unadjusted) · $0.0048/task · 2026-07-13
12 memorized-canon claims; naked BEATS harness here; grounding 0.83, hallucination 0.17 Test: exact McNemar on discordant pairs.
Run provenance
Report path: outputs/grounding/results.json
Task pack: grounding-canonical
Run: grounding
10.00Ground · tail
Compared with baselinen=14 · right DOI · split tail · Δ +9.29 vs Opus 4.8 · 13W 0L 1T on 14 paired · p<.001 (unadjusted) · $0.0046/task · 2026-07-13
14 niche 2024-25 claims; harness EARNS its keep here; grounding 1.00, hallucination 0.00 Test: exact McNemar on discordant pairs.
Run provenance
Report path: outputs/grounding-recent/results.json
Task pack: grounding-tail
Run: grounding-recent
—Find
Not yet tested.
—Extract
Not yet tested.
—Analyze
Not yet tested.
—Compute
Not yet tested.
—Synthesize
Not yet tested.
—Review
Not yet tested.
—Hypothesize
Not yet tested.