Reading Attention Patterns (first pass)
Pulled up attention patterns on Gemma with circuitsvis for a handful of prompts, no hypothesis going in. Mostly noise. What I could actually tell apart:
- a few heads: strict previous-token attention
- a few heads: strict first-token / attention-sink behavior
- most heads: nothing legible from eyeballing alone
Suspect this’ll make a lot more sense once patching tells me which heads matter for something specific, instead of going in blind.