▪ AI safety · alignment research unrulyabstractions.com
← unruly abstractions

Temporal Preference Concepts and their Functions in a Large Language Model

Unruly Abstractions May 11, 2026 Empirical AISC SPAR
Figure from Temporal Preference Concepts and their Functions in a Large Language Model

Abstract

Causally localizes a subgraph for temporal preference in a distilled LLM (Qwen3-4B-Instruct-2507) using gradient attribution and activation patching, with steering vectors as suggestive control