the interpretability gap
Picture two lines on a chart racing forward. One is how capable AI systems are becoming; the other is how well we can actually understand what is happening inside them. Today the capability line is sprinting and the understanding line is walking. The interpretability gap is the name for that distance: the difference between what interpretability can currently explain and what we would need to understand to make strong safety guarantees about powerful models.
Concretely, the gap shows up as a mismatch of scale and stakes. Current methods can reverse-engineer specific behaviors in small or medium models, extract large feature dictionaries, and verify isolated circuits — real progress. But what safety would ideally want is to read a frontier model's goals reliably, certify that it harbors no hidden plan to deceive, and do so faithfully and at full scale. Between 'we explained an induction circuit' and 'we can certify this system is not deceptively misaligned' lies an enormous, only-partly-charted distance, made worse by superposition, the sheer size of models, and the ever-present risk of interpretation illusions.
This gap is the honest backdrop to every interpretability claim, and people draw different conclusions from it. Optimists argue that the field is young and accelerating, that tools like sparse autoencoders are closing the gap fast, and that interpretability could become a cornerstone of safety. Skeptics worry that capabilities may keep outrunning understanding, that some of what we want to verify may be very hard or impossible to read out, and that we should not stake safety on a microscope we have not built yet. Where exactly the gap stands, and whether it is narrowing fast enough, is genuinely debated.
We can confidently say 'this attention head copies repeated patterns'. We cannot yet say 'this model has no hidden goal to mislead its operators'. The chasm between those two kinds of statement is the interpretability gap in one sentence.
The interpretability gap: the distance between what we can explain and what safety would require.
The size of the gap and how fast it is closing are contested. Treat interpretability as a maturing, promising tool to combine with other safety measures, not as a solved guarantee.