โ–ช AI safety ยท alignment research unrulyabstractions.com
โ† unruly abstractions

interpretability

Research on mechanistic interpretability, category-theoretic interpretability, activation patching, and circuit analysis in large language models.

  1. Measures how persona geometry shifts under the narrow fine-tuning that causes emergent misalignment, and steers against that shift to reduce misalignment across the Qwen family

  2. Causally localizes a subgraph for temporal preference in a distilled LLM (Qwen3-4B-Instruct-2507) using gradient attribution and activation patching, with steering vectors as suggestive control

  3. Which Circuit Is it?
    GroundlessLessWrong Mar 24, 2026

    A sequence on the foundations of interpretability through toy-model experiments, where subcircuits can be exhaustively cataloged and carefully modified

  4. Category-Theoretic Wanderings into Interpretability
    Queer in AI @ EurIPS Dec 5, 2025

    A technical autotheory exploration of AI interpretability through category theory