Reading sequence
AI from Scratch
The whole path in order: linear algebra intuition, tokenizers, embeddings, attention, the full transformer forward and backward pass, optimizers, and two complete builds (autograd in NumPy, GPT-2 in PyTorch). Each part assumes the one before it.
19 parts
312 min total
- Linear Algebra for Machine Learning: The Visual Intuition Linear algebra for machine learning explained visually: vectors, matrices, dot products, determinants, eigenvectors, and the basis changes that make it click.
- Backpropagation from Scratch: Build an Autograd Engine How does backpropagation work? Build a working autograd engine from scratch in ~80 lines of Python: computation graphs, chain rule, reverse-mode autodiff.
- Numerical Gradient Checking: Debug Your Autograd Engine What is numerical gradient checking? The central difference formula, why you never train with finite differences, and a grad checker that catches real bugs.
- Byte Pair Encoding (BPE) Explained: How GPT Tokenizers Work What is byte pair encoding (BPE)? How GPT tokenizers turn text into IDs: pretokenization, merge rules, byte-level vocab, and a from-scratch build.
- Token Embeddings Explained: How LLMs Turn IDs Into Vectors What is a token embedding? How LLMs map token IDs to learned vectors: the embedding matrix, gather vs one-hot matmul, and a from-scratch NumPy build.
- Positional Encoding Explained: How Transformers Learn Order What is positional encoding in transformers? Why attention is order-blind, learned vs sinusoidal embeddings, and a from-scratch NumPy implementation.
- Cross-Entropy Loss Explained: From Logits to LLM Training What is cross-entropy loss? Derive softmax and negative log-likelihood, work a 3-class example, and see why MSE fails for language model classification.
- Adam and AdamW Explained: How LLMs Update Their Weights How does the Adam optimizer work? Derive momentum, RMSprop and bias correction, see why AdamW decouples weight decay, and walk a numeric update by hand.
- How GPT Actually Works: A Visual Guide to Transformers How does a GPT actually work? A visual walkthrough of transformers: tokens, word embeddings, dot products, softmax and temperature, with real GPT-3 numbers.
- How Attention Works in Transformers: Queries, Keys, Values How does attention work in transformers? Queries, keys, values, the attention pattern, masking, and multi-head attention, with real GPT-3 parameter counts.
- Transformer from Scratch: Forward Pass and Backprop by Hand One full transformer training step worked by hand: embeddings, positional encoding, attention, layer norm, cross-entropy loss, backprop, and Adam.
- GPT Math Explained: The Full Forward Pass Beyond Attention How does GPT math work end to end? One training step from token IDs through Q/K/V, attention, MLP, logits, cross-entropy loss, backprop, and AdamW.
- Build GPT-2 from Scratch in PyTorch: A Full Walkthrough How to build GPT-2 from scratch in PyTorch: tokenization, causal self-attention, transformer blocks, weight tying, and a 124M training loop that runs.
- DeepSeek V4 Explained: Long-Context Engineering and Math A verification-first look at DeepSeek V4 claims, plus the real math behind sparse attention, KV-cache scaling, and long-context training stability.
- Build a Mini LLM from Scratch in NumPy: RoPE, GQA, SwiGLU How to build a mini LLM from scratch in NumPy: RoPE, GQA, QK-Norm, SwiGLU, tied embeddings, and a 3.87M chat companion with full architecture visuals.
- Natural Language Inference Explained: Entailment in NLP What is natural language inference (NLI)? How models label a premise and hypothesis as entailment, contradiction, or neutral on SNLI, MNLI, and BERT.
- Looped Transformers Explained: Recurrent Depth and Astra What is a looped transformer? How recurrent depth reuses layers, why Nanbeige 4.2 runs a 22-layer stack twice, and why that does not hide chain of thought.
- DeepSeek V4 Inside: One Token Through Every Block How does DeepSeek V4 process one token? I follow the byte p in deepseek through CSA, HCA, mHC, MoE, and MTP on a toy width, with Flash-0731 sizes.
- Convolutional Neural Networks: CNN Math Explained How do convolutional neural networks work? CNN math explained with kernels, padding, stride, pooling, receptive fields, and a full backpropagation example.