differential treatment
A model can treat two groups of people differently without explicit markers. We build audits that detect this differential treatment in deployed models and in models with secret loyalties.
meaningful diversity
Generative models amplify the biases in their training data and collapse toward typical outputs. We treat this homogenization as an AI safety problem and formalize what meaningful diversity requires.
circuits and concepts
We localize concepts inside language models and measure their function in behavior. Our evidence comes from activation patching, steering, and toy models small enough to catalog exhaustively.










