Currently the library doesn't support Attention Output (hook_z) SAEs. I personally use these a ton (and know of a few other groups working with them), and it would be great to just use sae_vis out of the box! I think this would be an easy change.
Relatedly, would be great to support DFA by source position for the hook_z dashboards, as this makes interpreting attention output features way easier. Example: Induction features are tricky to spot with max activating examples, but obvious with DFA.
Currently the library doesn't support Attention Output (hook_z) SAEs. I personally use these a ton (and know of a few other groups working with them), and it would be great to just use sae_vis out of the box! I think this would be an easy change.
Relatedly, would be great to support DFA by source position for the hook_z dashboards, as this makes interpreting attention output features way easier. Example: Induction features are tricky to spot with max activating examples, but obvious with DFA.