Coming July 30, 2026
A practical, evidence-based model for designing AI agents that check their own work — so reliability comes from the system, not from buying a bigger model.
Preregistered · Open · Free for everyone
per correct result, by switching to smaller models that self-verify
through structured, rule-based verification — no AI judges
success rate gain, even making small models outperform flagship AI
cheaper models × fewer wasted runs = a structurally smaller footprint