How to develop an unbiased ML model for loan allocations?

Machine Learning on AWS

A large retail bank wants to develop an ML system to help the risk management team decide on loan allocations for different demographics. What must the bank do to develop an unbiased ML model?

  1. Reduce the size of the training dataset.
  2. Ensure that the ML model predictions are consistent with historical results.
  3. Create a different ML model for each demographic group.
  4. Measure class imbalance on the training dataset. Adapt the training process accordingly. Source Reference Answer

Community Votes

D
100%

100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests the understanding of data preprocessing for fairness, specifically identifying that class imbalance is a primary source of bias in supervised learning tasks like loan approval.

Developing unbiased ML models requires measuring and addressing class imbalance in the training dataset to ensure fair representation across all demographic groups. Community consensus confirms that ignoring data distribution leads to biased predictions favoring majority classes.

Candidates often choose B (consistent with historical results), assuming that aligning with past decisions ensures fairness, but this actually perpetuates existing historical biases rather than correcting them.

Community Discussion (5 comments)

kopper2019 👍 1
D. Measure class imbalance on the training dataset. Adapt the training process accordingly.
Jessiii 👍 1 Selected: D
When developing an unbiased machine learning (ML) model, it's crucial to address issues like class imbalance in the training data. Class imbalance refers to the situation where certain classes (or demographic groups, in this case) are underrepresented compared to others. If class imbalance exists, the model might learn to favor the majority class and perform poorly on minority classes, leading to biased predictions.
Moon 👍 2 Selected: D
D. Measure class imbalance on the training dataset. Adapt the training process accordingly: This is the correct answer. Class imbalance occurs when one class (e.g., loan approval) is significantly more represented in the training data than another. This can lead to biased models that favor the majority class. Measuring and addressing class imbalance (e.g., through resampling or weighting techniques) is crucial for building fair models. Why not B? B. Ensure that the ML model predictions are consistent with historical results: If historical results reflect existing biases in lending practices, ensuring consistency with them will simply perpetuate those biases. This is the opposite of what is desired.
may2021_r 👍 1 Selected: D
The correct answer is D. Measuring and addressing class imbalance in training data is essential for developing unbiased ML models.
aws_Tamilan 👍 1 Selected: D
D. Measure class imbalance on the training dataset. Adapt the training process accordingly. Explanation: To develop an unbiased ML model, it is crucial to ensure that the training dataset represents all demographic groups fairly and that the model is not influenced by biases in the data.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option D is correct because class imbalance is a fundamental cause of bias in machine learning models. When one class (e.g., approved loans) vastly outnumbers another (e.g., denied loans or specific minority demographics), the model tends to optimize for accuracy by predicting the majority class, thereby failing to generalize for underrepresented groups. By measuring this imbalance and adapting the training process (via techniques like resampling, SMOTE, or class weighting), developers can create a more equitable model.

Why the Other Options Are Wrong

Option A is incorrect because reducing dataset size typically increases variance and reduces model performance, making it harder to detect or mitigate bias. Option B is dangerous because historical results often contain embedded societal or institutional biases; replicating them does not create an 'unbiased' model. Option C suggests creating separate models per demographic, which can lead to disparate impact if the models are not rigorously audited and may violate fairness constraints by treating groups differently rather than ensuring equal treatment.

Community Comment Notes

All provided comments unanimously support Option D. Users highlight that class imbalance directly causes models to favor the majority class, leading to poor performance on minority classes. One comment explicitly notes that ensuring fair representation is crucial for unbiased outcomes, reinforcing the link between data balance and model fairness.

Official Reference

Exam Strategy

When asked about 'fairness' or 'bias' in ML questions, always look for options related to data quality, such as checking for class imbalance, missing values, or skewed distributions. Avoid options that suggest using historical data as a ground truth, as history is often biased.

Related Analysis

Practice All AIF-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full AIF-C01 Practice Test →

← Back to AIF-C01 Study Guide