All topics

Attention

Every post tagged “Attention”, newest first.

4 posts
GPT 10 min

GPT Math Explained: The Full Forward Pass Beyond Attention

How does GPT math actually work end to end? Follow one training step from token IDs through Q/K/V, multi-head attention, MLP, logits, cross-entropy loss, backprop, and AdamW — past O = Attention(Q,K,V).

Transformers 20 min

Transformer from Scratch: The Full Forward Pass, Backprop, and Weight Update Math

One full transformer training step worked by hand — embeddings, positional encoding, attention, layer norm, the encoder-decoder, cross-entropy loss, backprop, and Adam.

DeepSeek 14 min

DeepSeek V4 Claims Examined: Long-Context Engineering, Math, and Verification

A verification-first analysis of circulating DeepSeek V4 claims, plus the real mathematics behind sparse attention, KV-cache scaling, stable residual mixing, and long-context training.

Attention 13 min

How Attention Works in Transformers: Queries, Keys and Values

How does attention work in transformers? A visual, step-by-step explanation of the attention mechanism — queries, keys, values, the attention pattern, masking, the low-rank value trick, and multi-head attention — with the real GPT-3 parameter counts.