Gruhesh Sri Sai Karthik Kurra

Notes on AI & ML

Gruhesh Sri Sai Karthik Kurra

Deep-dives into transformers, autograd, and the math behind modern LLMs — written while I build things from scratch.

Also on gruheshkurra.com

Latest posts

View all
Deep Learning 16 min

Convolutional Neural Networks: CNN Math Explained

How do convolutional neural networks work? CNN math explained with kernels, padding, stride, pooling, receptive fields, and a full backpropagation example.

DeepSeek 22 min

DeepSeek V4 Inside: One Token Through Every Block

How does DeepSeek V4 process one token? I follow the byte p in deepseek through CSA, HCA, mHC, MoE, and MTP on a toy width, with Flash-0731 sizes.

Transformers 14 min

Looped Transformers Explained: Recurrent Depth and Astra

What is a looped transformer? How recurrent depth reuses layers, why Nanbeige 4.2 runs a 22-layer stack twice, and why that does not hide chain of thought.

NLP 16 min

Natural Language Inference Explained: Entailment in NLP

What is natural language inference (NLI)? How models label a premise and hypothesis as entailment, contradiction, or neutral on SNLI, MNLI, and BERT.

LLM 11 min

Build a Mini LLM from Scratch in NumPy: RoPE, GQA, SwiGLU

How to build a mini LLM from scratch in NumPy: RoPE, GQA, QK-Norm, SwiGLU, tied embeddings, and a 3.87M chat companion with full architecture visuals.

GPT 10 min

GPT Math Explained: The Full Forward Pass Beyond Attention

How does GPT math work end to end? One training step from token IDs through Q/K/V, attention, MLP, logits, cross-entropy loss, backprop, and AdamW.

Deep Learning 11 min

Adam and AdamW Explained: How LLMs Update Their Weights

How does the Adam optimizer work? Derive momentum, RMSprop and bias correction, see why AdamW decouples weight decay, and walk a numeric update by hand.

Deep Learning 9 min

Cross-Entropy Loss Explained: From Logits to LLM Training

What is cross-entropy loss? Derive softmax and negative log-likelihood, work a 3-class example, and see why MSE fails for language model classification.

Tokenization 21 min

Byte Pair Encoding (BPE) Explained: How GPT Tokenizers Work

What is byte pair encoding (BPE)? How GPT tokenizers turn text into IDs: pretokenization, merge rules, byte-level vocab, and a from-scratch build.

Explore