Reading Attention Patterns (first pass)

Pulled up attention patterns on Gemma with circuitsvis for a handful of prompts, no hypothesis going in. Mostly noise. What I could actually tell apart:

  • a few heads: strict previous-token attention
  • a few heads: strict first-token / attention-sink behavior
  • most heads: nothing legible from eyeballing alone

Suspect this’ll make a lot more sense once patching tells me which heads matter for something specific, instead of going in blind.