Research on AI control: how deployed models behave differently toward user groups, and how to audit and measure that difference.
Detects secret loyalties in LLMs by measuring how a target model behaves differently around user groups that mention a candidate principal, read against the target's own base model
Builds a test harness that measures how accurately humans detect which of two assistants carries a secret loyalty, and compares them against automated baselines