▪ AI safety · alignment research unrulyabstractions.com

unruly abstractions

interpretability · evals · llm bias

An independent AI safety research.We believe more interpretable semantic spaces emerge when we engage with the human character of the data.

01

differential treatment

A model can treat two groups of people differently without explicit markers. We build audits that detect this differential treatment in deployed models and in models with secret loyalties.

  1. Can Humans Detect Secret Loyalties in LLMs?
    BlueDot Hackathon 2026Jul 2026
02

meaningful diversity

Generative models amplify the biases in their training data and collapse toward typical outputs. We treat this homogenization as an AI safety problem and formalize what meaningful diversity requires.

  1. The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
    Pluralistic Alignment @ ICMLCulture x AI @ ICMLJul 2026
03

circuits and concepts

We localize concepts inside language models and measure their function in behavior. Our evidence comes from activation patching, steering, and toy models small enough to catalog exhaustively.

  1. Which Circuit Is it?
    GroundlessLessWrongMar 2026