▪ AI safety · alignment research unrulyabstractions.com
← unruly abstractions

Discovering Instrumental Differential Treatment in Language Models

Ian Rios-Sialer, Eli Wang October 4, 2026 Theoretical + Empirical
Figure from Discovering Instrumental Differential Treatment in Language Models

Abstract

Formalizes instrumental differential treatment and gives a pipeline in which helper LLMs propose user groups, probing prompts and judges, then recover the groups a target model treats differently and the behaviors that separate them