Which Type of Bias Causes an ML Model to Disproportionately Flag a Specific Ethnic Group?
A company has installed a security camera. The company uses an ML model to evaluate the security camera footage for potential thefts. The company has discovered that the model disproportionately flags people who are members of a specific ethnic group. Which type of bias is affecting the model output?
Community Votes
85% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests the ability to distinguish between sampling bias and measurement bias by identifying whether the root cause lies in unrepresentative data selection versus flawed data collection or labeling processes.
This question tests knowledge of sampling bias in machine learning, where a model disproportionately flags a specific ethnic group due to unrepresentative training data. Community consensus strongly supports sampling bias as the correct answer, highlighting the importance of balanced datasets in ML fairness.
Candidates often choose measurement bias (Option A) because they associate the disproportionate flagging with flawed measurement processes, failing to recognize that the core issue is unrepresentative training data rather than how the data was measured or labeled.
Community Discussion (8 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Understanding the Correct Answer: Sampling Bias
Sampling bias occurs when the training data used for a machine learning model is not representative of the entire population or the range of scenarios the model will encounter in production. In this scenario, the security camera model disproportionately flags people from a specific ethnic group, which strongly suggests that the training dataset had an overrepresentation or underrepresentation of certain ethnic groups.
When training data lacks diversity, the model learns patterns that reflect the skewed distribution rather than the true population distribution. This leads to systematic errors where certain groups are unfairly targeted or overlooked.
Why the Other Options Are Incorrect
Measurement bias (Option A) refers to errors in how data is collected, labeled, or measured—not in the selection of the data itself. While some candidates chose this option, measurement bias would apply if the camera's sensors or the labeling process systematically misidentified certain ethnic groups, rather than if the training data itself was unrepresentative.
Observer bias (Option C) occurs when human annotators or observers introduce their own subjective biases into the data labeling process. This is a subset of measurement bias and doesn't directly apply to the scenario described.
Confirmation bias (Option D) is a cognitive bias where humans seek or interpret information in ways that confirm their preexisting beliefs. This is a human psychological phenomenon and doesn't directly apply to ML model training data issues.
Community Insights
The community overwhelmingly supports Option B (Sampling bias) with 85% of votes. As noted by multiple users, the key indicator is that the model's predictions are skewed due to unrepresentative training data, which is the hallmark of sampling bias. The minority who chose measurement bias focused on the flawed outcomes but missed the root cause: the training data selection process.
Official Reference
Exam Strategy
When identifying bias types in ML scenarios, focus on the root cause: if the issue stems from unrepresentative data selection, it's sampling bias; if it stems from flawed measurement or labeling processes, it's measurement bias. Always trace the problem back to its origin in the data pipeline.
Related Analysis
Practice All AIF-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full AIF-C01 Practice Test →