Transformer from Scratch: The Full Forward Pass, Backprop, and Weight Update Math
One full transformer training step worked by hand — embeddings, positional encoding, attention, layer norm, the encoder-decoder, cross-entropy loss, backprop, and Adam.