
Models & Stacks
Mixture of Experts Guide
How Mixtral and DeepSeek-V3 run 671B parameters while waking only 37B per token, and what that does to cost.
Published June 2026
Look inside
DeepSeek-V3 has 671B parameters but only 37B wake up per token. A router activates 8 of 256 experts plus 2 shared ones; the rest stay asleep. How Mixtral and DeepSeek actually route, what it does to inference cost, and the numbers the hype rounds off.
Or get the whole library
This is one of 154 guides in the Vault. Get every guide plus every future one with lifetime access for $297, less than buying 16 on their own.
Get the Vault →

