Transformer foundations, first pass

Starting properly today. Going breadth-first through the fundamentals section first: prerequisites, architecture intro, attention mechanism, MLPs, decoding strategies. Most of this is genuinely review — nothing here that a few years of DL work hasn’t already covered, so moved through it quickly rather than reading every word carefully.

Started implementing a transformer from scratch in PyTorch alongside the reading, mostly to make sure nothing’s rusty and to have something to poke at once the interpretability-specific stuff starts. Attention and MLP blocks straightforward. Left LayerNorm, QK/OV circuits, and composition/virtual heads for next time — skimmed them today and they’re clearly denser than the rest, want to actually sit with them properly instead of rushing.