
Models & Stacks
Self-Improving Model Setup Guide
The generate-verify-fine-tune loop that took DeepSeek-R1 from 15.6% to 77.9% on AIME with zero human labels.
Published June 2026
Look inside
Generate, verify, fine-tune, repeat. RLVR, STaR, and ReST took DeepSeek-R1 from 15.6% to 77.9% on AIME with zero human labels. It only works where outcomes are verifiable: the domains table shows where the loop compounds and where it collapses.
Or get the whole library
This is one of 154 guides in the Vault. Get every guide plus every future one with lifetime access for $297, less than buying 16 on their own.
Get the Vault →

