Reading Attention Patterns, take two

Went back to the same attention pattern visualizations on Gemma from over a week ago, this time only looking at heads path patching had flagged. Completely different experience — patterns that looked like noise before are obviously doing something specific once you know what to look for. One flagged head cleanly attends name-token → previous occurrence of that name, exactly the “copying” behavior the patching result implied. Reading attention patterns cold isn’t a great use of time; reading them with a hypothesis in hand is a different tool entirely.