โ–ช AI safety ยท alignment research unrulyabstractions.com
โ† unruly abstractions

all work

Every paper, proceeding, and research note, in chronological order.

  1. Steers Qwen3.6-27B to disidentify with being an AI assistant along contrastive behaviour directions, then shows the model cannot introspectively detect the steering

  2. Measures how persona geometry shifts under the narrow fine-tuning that causes emergent misalignment, and steers against that shift to reduce misalignment across the Qwen family

  3. Proposes a pipeline that discovers and measures how deployed LLMs treat queer and cisheterosexual users differently, without relying on explicit markers

  4. Detects secret loyalties in LLMs by measuring how a target model behaves differently around user groups that mention a candidate principal, read against the target's own base model

  5. The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
    Pluralistic Alignment @ ICMLCulture x AI @ ICML Jul 2026

    Foregrounds homogenization as a central AI safety concern; uses critical and queer theory to formalize diversity, normativity, and xeno-reproduction in LLMs

  6. Can Humans Detect Secret Loyalties in LLMs?
    BlueDot Hackathon 2026 Jul 2026

    Builds a test harness that measures how accurately humans detect which of two assistants carries a secret loyalty, and compares them against automated baselines

  7. Characterizes social bias in Spanish-prompted LLMs across model sizes with the SESGO benchmark; smaller models show more bias and yield more to prompt scaffolds

  8. Causally localizes a subgraph for temporal preference in a distilled LLM (Qwen3-4B-Instruct-2507) using gradient attribution and activation patching, with steering vectors as suggestive control

  9. Formalizing xeno-reproduction as structure-aware diversity pursuit to mitigate homogenization in generative AI

  10. Which Circuit Is it?
    GroundlessLessWrong Mar 24, 2026

    A sequence on the foundations of interpretability through toy-model experiments, where subcircuits can be exhaustively cataloged and carefully modified

  11. Category-Theoretic Wanderings into Interpretability
    Queer in AI @ EurIPS Dec 5, 2025

    A technical autotheory exploration of AI interpretability through category theory

  12. Xenoreproduction: AI Safety against Homogenization
    Queer in AI @ EurIPS Dec 5, 2025

    Exploring diversity and alternative modes of reproduction as an AI safety objective