Which Type of Bias Causes an ML Model to Disproportionately Flag a Specific Ethnic Group?

A company has installed a security camera. The company uses an ML model to evaluate the security camera footage for potential thefts. The company has discovered that the model disproportionately flags people who are members of a specific ethnic group. Which type of bias is affecting the model output?

  1. Measurement bias
  2. Sampling bias Source Reference Answer
  3. Observer bias
  4. Confirmation bias

Community Votes

B
85%
A
15%

85% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests the ability to distinguish between sampling bias and measurement bias by identifying whether the root cause lies in unrepresentative data selection versus flawed data collection or labeling processes.

This question tests knowledge of sampling bias in machine learning, where a model disproportionately flags a specific ethnic group due to unrepresentative training data. Community consensus strongly supports sampling bias as the correct answer, highlighting the importance of balanced datasets in ML fairness.

Candidates often choose measurement bias (Option A) because they associate the disproportionate flagging with flawed measurement processes, failing to recognize that the core issue is unrepresentative training data rather than how the data was measured or labeled.

Community Discussion (8 comments)

Willdoit 👍 2 Selected: B
Sampling bias occurs when the data used to train the machine learning model is not representative of the entire population or the range of possible scenarios. If the model disproportionately flags people from a specific ethnic group, it suggests that the data used to train the model may have an overrepresentation or underrepresentation of certain ethnic groups, leading to biased predictions.
Jessiii 👍 1 Selected: B
B. Sampling bias occurs when the data used to train the model is not representative of the entire population. In this case, if the training data contains an overrepresentation or underrepresentation of certain ethnic groups, the model may disproportionately flag individuals from specific ethnic groups, leading to biased outcomes.
pavankvv 👍 1 Selected: A
Measurement bias occurs when the data used to train the machine learning model contains inherent biases or inaccuracies, leading to biased outputs or predictions. In the given scenario, the security camera model is disproportionately flagging people from a specific ethnic group
Moon 👍 2 Selected: B
B: Sampling bias Explanation: Sampling bias occurs when the training data used for an ML model is not representative of the real-world population. In this case, the model disproportionately flags members of a specific ethnic group, likely because the training dataset was not balanced or representative of all groups. This leads to skewed predictions that unfairly target certain populations.
RightAnswers 👍 1 Selected: B
Sampling bias: occurs when the data used to train the ML model is not representative of the overall population, leading to the model performing poorly on certain groups, like in this case where the model is disproportionately flagging people from a specific ethnic group. Why the other options are not correct: Measurement bias: This refers to errors in the way data is collected or measured, which isn't directly related to the ethnic group bias in this scenario. Observer bias: This happens when a human observer's personal biases influence their interpretation of data, which isn't applicable here as the model is making the evaluations automatically. Confirmation bias: This refers to the tendency to seek out information that confirms existing beliefs, which isn't relevant to the training data used to develop the ML model.
eesa 👍 1 Selected: B
B. Sampling bias Explanation: Sampling bias occurs when the data used to train a model does not accurately represent the diversity of the population or real-world scenarios the model will encounter. In this case, if the training data for the security camera footage had an overrepresentation or underrepresentation of certain ethnic groups, the model may disproportionately flag members of that group as potential theft suspects. This leads to biased predictions due to imbalanced or unrepresentative training data.
aws4myself 👍 1 Selected: A
A. Measurement bias Measurement bias occurs when the measurement process itself is flawed, leading to systematic errors. In this case, the model is likely biased due to the way it's trained on data that may not be representative of the entire population. This can lead to the model incorrectly associating certain characteristics with criminal behavior, particularly for individuals from underrepresented groups.
Blair77 👍 4 Selected: B
B - Sampling bias occurs when the data used to train a model is not representative of the population or real-world scenarios it's meant to analyze. This leads to skewed results that favor or disfavor certain groups.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Understanding the Correct Answer: Sampling Bias

Sampling bias occurs when the training data used for a machine learning model is not representative of the entire population or the range of scenarios the model will encounter in production. In this scenario, the security camera model disproportionately flags people from a specific ethnic group, which strongly suggests that the training dataset had an overrepresentation or underrepresentation of certain ethnic groups.

When training data lacks diversity, the model learns patterns that reflect the skewed distribution rather than the true population distribution. This leads to systematic errors where certain groups are unfairly targeted or overlooked.

Why the Other Options Are Incorrect

Measurement bias (Option A) refers to errors in how data is collected, labeled, or measured—not in the selection of the data itself. While some candidates chose this option, measurement bias would apply if the camera's sensors or the labeling process systematically misidentified certain ethnic groups, rather than if the training data itself was unrepresentative.

Observer bias (Option C) occurs when human annotators or observers introduce their own subjective biases into the data labeling process. This is a subset of measurement bias and doesn't directly apply to the scenario described.

Confirmation bias (Option D) is a cognitive bias where humans seek or interpret information in ways that confirm their preexisting beliefs. This is a human psychological phenomenon and doesn't directly apply to ML model training data issues.

Community Insights

The community overwhelmingly supports Option B (Sampling bias) with 85% of votes. As noted by multiple users, the key indicator is that the model's predictions are skewed due to unrepresentative training data, which is the hallmark of sampling bias. The minority who chose measurement bias focused on the flawed outcomes but missed the root cause: the training data selection process.

Official Reference

Exam Strategy

When identifying bias types in ML scenarios, focus on the root cause: if the issue stems from unrepresentative data selection, it's sampling bias; if it stems from flawed measurement or labeling processes, it's measurement bias. Always trace the problem back to its origin in the data pipeline.

Related Analysis

Practice All AIF-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full AIF-C01 Practice Test →

← Back to AIF-C01 Study Guide