Showing posts with label 3blue1brown. Show all posts
Showing posts with label 3blue1brown. Show all posts

Thursday, July 16, 2026

But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.
Job opportunities aligned to this audience: https://3b1b.co/talent
Early views and other perks for supporters: https://3b1b.co/support
Home page: https://www.3blue1brown.com

Manim animations by Aaron Gostein and Grant Sanderson
NanoGPT animation by Clayton Rabideau
3d black-box model by Paul Dancstep
Music by Vince Rubinetti

Timestamps

0:00 - Language trees and zipping
3:02 - Recap optimal codes
5:20 - Defining cross-entropy
8:26 - Intuition and examples
12:59 - Application to language trees
14:55 - Pre-training LLMs
20:38 - What makes this loss function best?
26:13 - Distillation
30:12 - 3b1b Talent
31:35 - KL Divergence 

* * * * *

Friday, October 4, 2024

How might LLMs store facts? [Multilayer Perceptrons, MLP]

Time stamps:

0:00 - Where facts in LLMs live
2:15 - Quick refresher on transformers
4:39 - Assumptions for our toy example
6:07 - Inside a multilayer perceptron
15:38 - Counting parameters
17:04 - Superposition
21:37 - Up next

Preceding videos in this series:

Sunday, April 7, 2024

Visualizing Attention: A Transformer's Heart

This is the second of three videos from 3Blue1Brown about how transformers work. Here's the first.

Timestamps:
0:00 - Recap on embeddings:
1:39 - Motivating examples:
4:29 - The attention pattern:
11:08 - Masking:
12:42 - Context size:
13:10 - Values:
15:44 - Counting parameters:
18:21 - Cross-attention:
19:19 - Multiple heads:
22:16 - The output matrix:
23:19 - Going deeper:
24:54 - Ending

You might also want to look at this post, where I have three videos where 3Blue1Brown visualizes the use of Fourier series to trace out fairly elaborate line drawings.

Tuesday, April 2, 2024

But what is a GPT? Visual intro to Transformers

This is the first of three videos from 3Blue1Brown about how transformers work. The other two will be released in the coming weeks, though if you’re a member of the Patreon channel you can review and comment on a draft of the second video in the series.

Timestamps

0:00 - Predict, sample, repeat
3:03 - Inside a transformer
6:36 - Chapter layout
7:20 - The premise of Deep Learning
12:27 - Word embeddings
18:25 - Embeddings beyond words
20:22 - Unembedding
22:22 - Softmax with temperature
26:03 - Up next

In an ideal world, a video like this would have been released much earlier, no later than the release of ChatGPT to the general public on November 30, 2022. That way those curious about how these creatures work, but not (quite) having the intellectual skills required to read the technical literature, would have had something to work with. That wouldn’t, however, have prevented the “stochastic parrots” nonsense, that’s more an ideological judgment as an intellectual. But it might have helped to lower the level of confusion attendant upon the release of ChatGPT.

As it is, it’s taken 488 days for this video to become available, and the series is not yet complete. Take that lag as an index of the problems we’re having adjusting to the implications of this technology. The lag isn’t anyone’s fault, not OpenAI’s, not 3Blue1Brown’s, no one’s. It’s just a property of the techno-socio-economic-cultural system in which we live.