Every paper, proceeding, and research note, in chronological order.
Measures how persona geometry shifts under the narrow fine-tuning that causes emergent misalignment, and steers against that shift to reduce misalignment across the Qwen family
Proposes a pipeline that discovers and measures how deployed LLMs treat queer and cisheterosexual users differently, without relying on explicit markers
Detects secret loyalties in LLMs by measuring how a target model behaves differently around user groups that mention a candidate principal, read against the target's own base model
Foregrounds homogenization as a central AI safety concern; uses critical and queer theory to formalize diversity, normativity, and xeno-reproduction in LLMs
Builds a test harness that measures how accurately humans detect which of two assistants carries a secret loyalty, and compares them against automated baselines
Characterizes social bias in Spanish-prompted LLMs across model sizes with the SESGO benchmark; smaller models show more bias and yield more to prompt scaffolds
Causally localizes a subgraph for temporal preference in a distilled LLM (Qwen3-4B-Instruct-2507) using gradient attribution and activation patching, with steering vectors as suggestive control
Formalizing xeno-reproduction as structure-aware diversity pursuit to mitigate homogenization in generative AI
A sequence on the foundations of interpretability through toy-model experiments, where subcircuits can be exhaustively cataloged and carefully modified
A technical autotheory exploration of AI interpretability through category theory
Exploring diversity and alternative modes of reproduction as an AI safety objective