Question
A company routes 80% of traffic to a champion model and 20% to a challenger. The aggregate weekly accuracy appears unchanged, but customer complaints started after the challenger rollout. The inference table includes `model_id`, predictions, labels, and timestamps. What analysis is most appropriate?