Concept Drift in Customer Data Causing F1 Score Drop After Months of Stable Model Performance

Answer Correct answer: A — Concept drift shifted the customer data distribution, so a model calibrated on the old baseline underperforms and Model Monitor reports an F1 deviation.

A company has deployed an XGBoost prediction model in production to predict if a customer is likely to cancel a subscription. The company uses Amazon SageMaker Model Monitor to detect deviations in the F1 score. During a baseline analysis of model quality, the company recorded a threshold for the F1 score. After several months of no change, the model's F1 score decreases significantly. What could be the reason for the reduced F1 score?

  1. Concept drift occurred in the underlying customer data that was used for predictions. Correct Answer
  2. The model was not sufficiently complex to capture all the patterns in the original baseline data.
  3. The original baseline data had a data quality issue of missing values.
  4. Incorrect ground truth labels were provided to Model Monitor during the calculation of the baseline.

Community Votes

A
100%

100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Concept drift means the statistical properties of the data the model scores have changed over time, so a model calibrated against an old distribution degrades on current data. The several-months-of-stability timeline rules out any cause that would have been visible at baseline time.

An XGBoost model predicts subscription cancellation and is monitored by SageMaker Model Monitor, which detects deviations in F1 score against a recorded baseline threshold. The model holds steady for several months and then the F1 score drops significantly, so the cause has to be something that changed in the world after the baseline was established.

Selecting a model complexity problem or a baseline data quality problem, both of which contradict the timeline. Insufficient model complexity would have depressed the F1 score at baseline, and missing values in the baseline data would have surfaced during the baseline analysis rather than months later.

Community Discussion (4 comments)

ninomfr64 👍 1 Selected: A
A. Yes, concept drift is an evolution of data that invalidates the data model. It happens when the statistical properties of the target variable, which the model is trying to predict, change over time in unforeseen ways. This causes problems because the predictions become less accurate as time passes. B. No, if it was the case the F1 would have been low since the begin and this is not justifying a change after months C. No, same as B D. No, incorrect labels in the baseline calculations would undermine F1 baseline value, but this is not explain a significant drop after months
motk123 👍 3 Selected: A
Concept Drift: Occurs when the statistical properties of the data used for predictions change over time, causing the model to underperform on current data. Why Not the Other Options? B. If the model complexity was insufficient, the issue would have been detected during the initial evaluation or baseline analysis, not after months of stable performance. C. A data quality issue would have impacted the model's performance immediately after deployment, not months later. D. Incorrect labels during baseline calculation could result in an inaccurate baseline F1 score, but it wouldn't explain a significant drop after stable performance over months.
Saransundar 👍 3 Selected: A
Concept Drift: Refers to the change in the statistical properties of the underlying data distribution over time --> Decrease F1 score --> perform poorly on new data
GiorgioGss 👍 2 Selected: A
Option A could be the only one possible reason for drifting "after several months".

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

The decisive detail is the timeline: the F1 score was stable for several months and then fell significantly. A change in the statistical properties of the underlying customer data over time is concept drift, and a model whose F1 threshold was calibrated on the old distribution will underperform once the incoming data has shifted, which is exactly the observed pattern. Model Monitor exists to surface precisely this kind of degradation, and the F1 deviation it reported is the symptom of the distribution shift rather than a modeling defect. The vote was unanimous at 100 for A. motk123 defined concept drift as a change in the statistical properties of the prediction data causing underperformance, and Saransundar tied it directly to the F1 decrease on new data.

Why the Other Options Are Wrong

The model not being complex enough to capture patterns in the baseline data (B) would have produced a low F1 score from the start, since the baseline analysis measured exactly that, so it cannot explain several months of acceptable performance followed by a decline. A data quality issue involving missing values in the original baseline data (C) would also have been detected during the baseline analysis itself, because Model Monitor evaluates the baseline data as part of establishing the threshold. Incorrect ground truth labels supplied during the baseline calculation (D) would bias the baseline threshold from the outset, again producing a wrong threshold at the start rather than a correct threshold that later degrades.

Community Comment Notes

The community was unanimous at 100 for A, and the comments converge on the timeline as the deciding evidence. GiorgioGss noted that concept drift is the only option consistent with the problem appearing after several months. motk123 and ninomfr64 both explained why the alternatives fail on timing grounds, with ninomfr64 pointing out that insufficient model complexity would have kept the F1 score low from the beginning. The unanimity here reflects a question where the distractors are ruled out by the stability timeline rather than by subtle technical distinctions.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide