A machine learning team discovered that fine-tuning a verifier model on customer-support traffic caused active harm due to distribution shift: the gray-zone positive rate drifted 5x over time, causing the model to approve wrong answers. They deployed three monitoring detectors—a z-test, a Page-Hinkley change-point detector, and a Kolmogorov-Smirnov test—to catch label-based drift and score distribution shifts automatically.