← LibraryField Notes
Self-Improving Model Setup Guide
Models & Stacks

Self-Improving Model Setup Guide

The generate-verify-fine-tune loop that took DeepSeek-R1 from 15.6% to 77.9% on AIME with zero human labels.

Published June 2026

Look inside

Generate, verify, fine-tune, repeat. RLVR, STaR, and ReST took DeepSeek-R1 from 15.6% to 77.9% on AIME with zero human labels. It only works where outcomes are verifiable: the domains table shows where the loop compounds and where it collapses.

$19Buy this guideinstant download

Or get the whole library

This is one of 154 guides in the Vault. Get every guide plus every future one with lifetime access for $297, less than buying 16 on their own.

Get the Vault →

More in Models & Stacks