Deep Learning 16 min

Convolutional Neural Networks: CNN Math Explained

How do convolutional neural networks work? CNN math explained with kernels, padding, stride, pooling, receptive fields, and a full backpropagation example.

DeepSeek 22 min

DeepSeek V4 Inside: One Token Through Every Block

How does DeepSeek V4 process one token? I follow the byte p in deepseek through CSA, HCA, mHC, MoE, and MTP on a toy width, with Flash-0731 sizes.

Transformers 14 min

Looped Transformers Explained: Recurrent Depth and Astra

What is a looped transformer? How recurrent depth reuses layers, why Nanbeige 4.2 runs a 22-layer stack twice, and why that does not hide chain of thought.

NLP 16 min

Natural Language Inference Explained: Entailment in NLP

What is natural language inference (NLI)? How models label a premise and hypothesis as entailment, contradiction, or neutral on SNLI, MNLI, and BERT.

LLM 11 min

Build a Mini LLM from Scratch in NumPy: RoPE, GQA, SwiGLU

How to build a mini LLM from scratch in NumPy: RoPE, GQA, QK-Norm, SwiGLU, tied embeddings, and a 3.87M chat companion with full architecture visuals.

GPT 10 min

GPT Math Explained: The Full Forward Pass Beyond Attention

How does GPT math work end to end? One training step from token IDs through Q/K/V, attention, MLP, logits, cross-entropy loss, backprop, and AdamW.

Deep Learning 9 min

Cross-Entropy Loss Explained: From Logits to LLM Training

What is cross-entropy loss? Derive softmax and negative log-likelihood, work a 3-class example, and see why MSE fails for language model classification.

Deep Learning 11 min

Adam and AdamW Explained: How LLMs Update Their Weights

How does the Adam optimizer work? Derive momentum, RMSprop and bias correction, see why AdamW decouples weight decay, and walk a numeric update by hand.

Embeddings 15 min

Token Embeddings Explained: How LLMs Turn IDs Into Vectors

What is a token embedding? How LLMs map token IDs to learned vectors: the embedding matrix, gather vs one-hot matmul, and a from-scratch NumPy build.

Positional Encoding 12 min

Positional Encoding Explained: How Transformers Learn Order

What is positional encoding in transformers? Why attention is order-blind, learned vs sinusoidal embeddings, and a from-scratch NumPy implementation.

Tokenization 21 min

Byte Pair Encoding (BPE) Explained: How GPT Tokenizers Work

What is byte pair encoding (BPE)? How GPT tokenizers turn text into IDs: pretokenization, merge rules, byte-level vocab, and a from-scratch build.

Deep Learning 14 min

Numerical Gradient Checking: Debug Your Autograd Engine

What is numerical gradient checking? The central difference formula, why you never train with finite differences, and a grad checker that catches real bugs.

Transformers 20 min

Transformer from Scratch: Forward Pass and Backprop by Hand

One full transformer training step worked by hand: embeddings, positional encoding, attention, layer norm, cross-entropy loss, backprop, and Adam.

DeepSeek 14 min

DeepSeek V4 Explained: Long-Context Engineering and Math

A verification-first look at DeepSeek V4 claims, plus the real math behind sparse attention, KV-cache scaling, and long-context training stability.

Attention 13 min

How Attention Works in Transformers: Queries, Keys, Values

How does attention work in transformers? Queries, keys, values, the attention pattern, masking, and multi-head attention, with real GPT-3 parameter counts.

Transformers 11 min

How GPT Actually Works: A Visual Guide to Transformers

How does a GPT actually work? A visual walkthrough of transformers: tokens, word embeddings, dot products, softmax and temperature, with real GPT-3 numbers.

Linear Algebra 12 min

Linear Algebra for Machine Learning: The Visual Intuition

Linear algebra for machine learning explained visually: vectors, matrices, dot products, determinants, eigenvectors, and the basis changes that make it click.

GPT-2 40 min

Build GPT-2 from Scratch in PyTorch: A Full Walkthrough

How to build GPT-2 from scratch in PyTorch: tokenization, causal self-attention, transformer blocks, weight tying, and a 124M training loop that runs.

Deep Learning 31 min

Backpropagation from Scratch: Build an Autograd Engine

How does backpropagation work? Build a working autograd engine from scratch in ~80 lines of Python: computation graphs, chain rule, reverse-mode autodiff.