Tracking my mechanistic interpretability self-study

less than 1 minute read

Since late May I've been working through mechanistic interpretability and AI alignment on my own, mostly following learnmechinterp.com's curriculum plus whatever papers, lectures, and courses came up along the way. Logging it here as I go.

It starts with why I picked mech interp over other parts of alignment, works through transformer foundations, a BlueDot AI Safety Fundamentals intensive I did in the middle of it, and then into the core interpretability curriculum — logit lens, activation patching, induction heads, probing, and onward. It’s still going.

Read the full log →