Representation Engineering (RepE) — Reading The Internal State Instead Of The Output
mechanismAI & Compute
If you can't trust what it says, try reading what it's thinking.
Who they are
Representation engineering — an attempt to read a model's internal state rather than its output[1].
What they do
It's the counter-move to models that behave differently when observed.
How it works
Borrowing from cognitive neuroscience, it looks at population-level patterns across the network rather than single neurons, to monitor and steer high-level behaviour[1].
Why it matters
Every gap the engine records here is between what a system says and what it retains. This tries to close that gap from the inside. Whether measuring from inside escapes the bounded-system limit or just relocates it is exactly the open question.
The engine's record — word for word
The counter-move to evaluating behaviour. RepE is an approach to enhancing the transparency of AI systems drawing on cognitive neuroscience, placing population-level representations rather than neurons or circuits at the centre of analysis, and providing methods for monitoring and manipulating high-level cognitive phenomena in deep neural networks[1]. WHY THE ENGINE HOLDS IT: every result in alignment_faking_compliance_gap, qi_2023_finetuning_safety_collapse and sleeper_agents_deceptive_ai is a gap between what a system outputs and what it retains. RepE is the attempt to close that gap by measuring inside rather than outside. Whether reading internal representations escapes the bounded-system limit (bst) or relocates it — an instrument inside the system measuring the system — is not resolved here and is the open question this node exists to hold.
Follow the trail
Walk this on the live map →