← All series
Model Distillation
A 7-part series: why frontier models are too expensive to deploy, the math that transfers knowledge into small ones, a hands-on build, and how to keep a distilled model good in production.

- 017 min
Why Everyone Runs a Model They Can't Afford
The problem nobody budgets for
- 027 min
Soft Labels: The Secret Ingredient
Why wrong answers teach more than right ones
- 037 min
The Three Flavors of Distillation: Response, Feature, and Relation
Copy the answers, copy the thinking, or copy the worldview
- 047 min
The Math That Makes It Work: KL Divergence and Temperature, Demystified
One loss function, two teachers — and the scaling factor everyone forgets
- 055 min
Hands-On: Distilling a Real Model (And Watching It Actually Work)
Theory meets a GPU — shrink a model, keep the accuracy, publish the receipts
- 066 min
Distillation in the Wild: DistilBERT, Mini Models, and Synthetic Data
How the pros actually do it — and the gray zone nobody likes to discuss
- 077 min
Production Distillation: Pipelines, Pitfalls, and What's Next
Distilling once is a project. Staying distilled is a discipline.





