Concept Drift in Customer Data Causing F1 Score Drop After Months of Stable Model Performance
A company has deployed an XGBoost prediction model in production to predict if a customer is likely to cancel a subscription. The company uses Amazon SageMaker Model Monitor to detect deviations in the F1 score. During a baseline analysis of model quality, the company recorded a threshold for the F1 score. After several months of no change, the model's F1 score decreases significantly. What could be the reason for the reduced F1 score?
Community Votes
100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
Concept drift means the statistical properties of the data the model scores have changed over time, so a model calibrated against an old distribution degrades on current data. The several-months-of-stability timeline rules out any cause that would have been visible at baseline time.
An XGBoost model predicts subscription cancellation and is monitored by SageMaker Model Monitor, which detects deviations in F1 score against a recorded baseline threshold. The model holds steady for several months and then the F1 score drops significantly, so the cause has to be something that changed in the world after the baseline was established.
Selecting a model complexity problem or a baseline data quality problem, both of which contradict the timeline. Insufficient model complexity would have depressed the F1 score at baseline, and missing values in the baseline data would have surfaced during the baseline analysis rather than months later.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The decisive detail is the timeline: the F1 score was stable for several months and then fell significantly. A change in the statistical properties of the underlying customer data over time is concept drift, and a model whose F1 threshold was calibrated on the old distribution will underperform once the incoming data has shifted, which is exactly the observed pattern. Model Monitor exists to surface precisely this kind of degradation, and the F1 deviation it reported is the symptom of the distribution shift rather than a modeling defect. The vote was unanimous at 100 for A. motk123 defined concept drift as a change in the statistical properties of the prediction data causing underperformance, and Saransundar tied it directly to the F1 decrease on new data.Why the Other Options Are Wrong
The model not being complex enough to capture patterns in the baseline data (B) would have produced a low F1 score from the start, since the baseline analysis measured exactly that, so it cannot explain several months of acceptable performance followed by a decline. A data quality issue involving missing values in the original baseline data (C) would also have been detected during the baseline analysis itself, because Model Monitor evaluates the baseline data as part of establishing the threshold. Incorrect ground truth labels supplied during the baseline calculation (D) would bias the baseline threshold from the outset, again producing a wrong threshold at the start rather than a correct threshold that later degrades.Community Comment Notes
The community was unanimous at 100 for A, and the comments converge on the timeline as the deciding evidence. GiorgioGss noted that concept drift is the only option consistent with the problem appearing after several months. motk123 and ninomfr64 both explained why the alternatives fail on timing grounds, with ninomfr64 pointing out that insufficient model complexity would have kept the F1 score low from the beginning. The unanimity here reflects a question where the distractors are ruled out by the stability timeline rather than by subtle technical distinctions.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →