Skill comparisons
All runs ↗| Task | AI alone | + Checklist | + Skill |
|---|---|---|---|
| Check a claim ↗n=8 · verdict accuracy | 8/8$0.0077 · 3.4s | 8/8$0.0059 · 3.1s | 8/8$0.0076 · 3.2s |
| Choose papers ↗n=4 · recall@5 · ceiling 50% | 43.8%$0.0302 · 3.1s | 43.8%$0.0314 · 4.0s | 43.8%$0.0336 · 3.6s |
| Extract information ↗n=8 · F1 · local procedure | 0.240$0.0124 · 7.4s | 0.306$0.0131 · 7.5s | 0.334$0.0165 · 10.3s |
| ↳ Wrappers removedPost-hoc rescore · same 8 answers | 0.400 | 0.341 | 0.391 |
| Analyze data ↗n=8 · 1 baseline tool-call failure | 7/8$0.0290 · 20.3s | 8/8$0.0166 · 13.1s | 8/8$0.0277 · 17.4s |
Score, then mean $ and seconds per attempt. One attempt per setup per task. Public development cases.
Analysis reproduction
1 case · 3 runs · $0.39| Setup | Correct | Replay + figure | Skill opened | Cost | Elapsed |
|---|---|---|---|---|---|
| AI alone | 3/3 | ✓ | — | $0.1162 | 131.3s |
| With checklist | 3/3 | ✓ | — | $0.1607 | 258.9s |
| Skill installed | 3/3 | ✓ | No | $0.1163 | 137.1s |
One supplied-code case. Installed skill never opened. Elapsed time includes file searches.
Source & run details
Reproduce a published structural-equation analysis ↗ · 2026-09-15 · skill revision 330c8e7.
One supplied-code teaching case with one attempt per setup. Three values come from one model fit; they are not three independent cases. The checklist includes a task-specific residual-variance reminder. All setups retain the native host’s built-in skills. Measured benchmark model calls only; local compute, runtime setup, storage and assistant development are excluded.
Methodology & limitations
Same model, inputs and limits within each comparison. Failures remain in the denominator. Pilot skills are supplied instructions; the native run tests an installed package. Counts are tasks, not necessarily independent datasets. No overall score.
Claims: verdict agreement with dataset labels. Papers: recall@5 from a fixed list. Extraction: entity/relation F1; malformed output scores zero. Analysis: code must pass original and changed-input checks. The wrapper rescore is a post-hoc sensitivity check with unchanged entity/relation content.
Model calls: pilot table $1.4741; full campaign including calibration and diagnostics $2.8401. Matched results ↗ · Format diagnostic ↗
Next experiments
- Explicit skill loading, scored as a separate diagnostic.
- Eight new datasets requiring method choice.
- Freeze the protocol; confirm on reserved sources.