Where the loss function for training LLMs comes from.
Job opportunities aligned to this audience: https://3b1b.co/talent
Early views and other perks for supporters: https://3b1b.co/support
Home page: https://www.3blue1brown.com
Manim animations by Aaron Gostein and Grant Sanderson
NanoGPT animation by Clayton Rabideau
3d black-box model by Paul Dancstep
Music by Vince Rubinetti
Timestamps
0:00 - Language trees and zipping
3:02 - Recap optimal codes
5:20 - Defining cross-entropy
8:26 - Intuition and examples
12:59 - Application to language trees
14:55 - Pre-training LLMs
20:38 - What makes this loss function best?
26:13 - Distillation
30:12 - 3b1b Talent
31:35 - KL Divergence
* * * * *
As someone who's been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is the closest thing to an ML PhD qualifying exam I've ever seen released publicly for free.
— fintex (@_yusufknl) July 19, 2026
Everyone thinks language models predict the next word. They don't. They… pic.twitter.com/KcdvO620Lw