โ–ช AI safety ยท alignment research unrulyabstractions.com
โ† unruly abstractions

ai control

Research on AI control: how deployed models behave differently toward user groups, and how to audit and measure that difference.

  1. Detects secret loyalties in LLMs by measuring how a target model behaves differently around user groups that mention a candidate principal, read against the target's own base model

  2. Can Humans Detect Secret Loyalties in LLMs?
    BlueDot Hackathon 2026 Jul 2026

    Builds a test harness that measures how accurately humans detect which of two assistants carries a secret loyalty, and compares them against automated baselines