▪ AI safety · alignment research unrulyabstractions.com
← unruly abstractions

Secret Loyalties as Instrumental Differential Treatment

Unruly Abstractions July 26, 2026 Empirical Apart Research
Figure from Secret Loyalties as Instrumental Differential Treatment

Abstract

Detects secret loyalties in LLMs by measuring how a target model behaves differently around user groups that mention a candidate principal, read against the target's own base model