โ–ช AI safety ยท alignment research unrulyabstractions.com
โ† unruly abstractions

all work

Every paper, proceeding, and research note, in chronological order.

  1. Measures how persona geometry shifts under the narrow fine-tuning that causes emergent misalignment, and steers against that shift to reduce misalignment across the Qwen family

  2. Proposes a pipeline that discovers and measures how deployed LLMs treat queer and cisheterosexual users differently, without relying on explicit markers

  3. Detects secret loyalties in LLMs by measuring how a target model behaves differently around user groups that mention a candidate principal, read against the target's own base model

  4. The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
    Pluralistic Alignment @ ICMLCulture x AI @ ICML Jul 2026

    Foregrounds homogenization as a central AI safety concern; uses critical and queer theory to formalize diversity, normativity, and xeno-reproduction in LLMs

  5. Can Humans Detect Secret Loyalties in LLMs?
    BlueDot Hackathon 2026 Jul 2026

    Builds a test harness that measures how accurately humans detect which of two assistants carries a secret loyalty, and compares them against automated baselines

  6. Characterizes social bias in Spanish-prompted LLMs across model sizes with the SESGO benchmark; smaller models show more bias and yield more to prompt scaffolds

  7. Causally localizes a subgraph for temporal preference in a distilled LLM (Qwen3-4B-Instruct-2507) using gradient attribution and activation patching, with steering vectors as suggestive control

  8. Formalizing xeno-reproduction as structure-aware diversity pursuit to mitigate homogenization in generative AI

  9. Which Circuit Is it?
    GroundlessLessWrong Mar 24, 2026

    A sequence on the foundations of interpretability through toy-model experiments, where subcircuits can be exhaustively cataloged and carefully modified

  10. Category-Theoretic Wanderings into Interpretability
    Queer in AI @ EurIPS Dec 5, 2025

    A technical autotheory exploration of AI interpretability through category theory

  11. Xenoreproduction: AI Safety against Homogenization
    Queer in AI @ EurIPS Dec 5, 2025

    Exploring diversity and alternative modes of reproduction as an AI safety objective