representation engineering (RepE)
Representation engineering is a top-down approach to interpretability and control that treats high-level concepts, honesty, harmlessness, power-seeking, emotion, as the primary unit of analysis, rather than starting from individual neurons or circuits. The stance is that complex behaviors are encoded in distributed population activity, so the productive move is to locate the direction in representation space that tracks a concept and then read it as a monitor or write it as a control.
Reading is done with representation reading: elicit the concept with stimulus pairs, collect activations, and use a linear or dimensionality-reduction method to extract a concept direction or a probe that detects when the concept is active. Controlling is done with representation control: perturb activations along that direction, or fine-tune with a loss defined in representation space, to amplify or suppress the concept. Activation steering is one concrete control instance under this umbrella.
RepE is appealing because it operates at the level humans care about and yields both transparency, a lie or sycophancy detector reading the residual stream, and steering, suppressing a harmful concept, from the same direction. The contrast with bottom-up mechanistic interpretability is deliberate; the cost is rigor, since a linear concept direction may be an approximation, can entangle correlated concepts, and a representation that predicts a concept is not proof the model uses it causally without intervention tests.