Tracking exactly what happens to the internal "circuitry" (induction heads) of a 2-layer attention-only Transformer when forced to undergo domain adaptation from prose to structured Python code.
research transformers pytorch ai-safety interpretability fine-tuning phase-transitions mechanistic-interpretability activation-patching transformer-circuits circuit-discovery induction-heads circuit-stability
-
Updated
Jul 15, 2026 - Python