Which algorithm groups customers by demographics and buying patterns?
A company wants to find groups for its customers based on the customers’ demographics and buying patterns. Which algorithm should the company use to meet this requirement?
Community Votes
80% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
This question tests the ability to distinguish between supervised learning (classification) and unsupervised learning (clustering), with the keyword 'find groups' being the critical indicator of a clustering task.
The question tests knowledge of unsupervised clustering algorithms, specifically K-means, for customer segmentation based on demographics and buying patterns. Community consensus strongly supports K-means as the correct choice for grouping unlabeled data into natural clusters.
Many candidates incorrectly choose K-nearest neighbors (k-NN) because it also contains the letter 'K' and involves grouping concepts, but k-NN is a supervised classification algorithm that requires labeled training data, not an unsupervised clustering method.
Community Discussion (6 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Understanding the Problem
The key phrase in this question is "find groups" for customers based on their characteristics. This immediately signals an unsupervised learning task, specifically clustering, where the goal is to discover natural groupings in data without pre-existing labels.
Why K-means is Correct
K-means is a classic clustering algorithm that partitions data into K distinct groups (clusters) based on feature similarity. In this scenario:
- Customer demographics (age, income, location) and buying patterns (purchase frequency, average spend, product categories) serve as the features
- K-means automatically discovers natural customer segments by minimizing the distance between data points within each cluster
- The result is distinct customer groups that can be used for targeted marketing and personalized offerings
Why Other Options Are Wrong
A. K-nearest neighbors (k-NN): This is a supervised classification algorithm that requires labeled training data. It classifies new data points based on the majority class of their K nearest neighbors. Since the company wants to "find groups" (discover unknown segments), not classify customers into known categories, k-NN is inappropriate.
C. Decision tree: This is also a supervised learning algorithm used for classification or regression. It requires labeled data to learn decision rules. Decision trees don't discover natural groupings; they classify based on learned patterns from labeled examples.
D. Support vector machine: SVM is another supervised classification algorithm that finds optimal boundaries between predefined classes. Like decision trees, it requires labeled training data and cannot discover unknown customer segments.
Community Insights
As noted by community members, the distinction is clear: "K-means - clustering - groups" while the other algorithms are for classification tasks. The keyword "groups" is the giveaway that this is a clustering problem requiring an unsupervised approach.
Official Reference
Exam Strategy
When you see keywords like 'find groups,' 'segment,' or 'discover patterns' without mention of labels or categories, immediately think 'unsupervised learning' and 'clustering.' Eliminate all supervised algorithms (k-NN, decision trees, SVM, logistic regression) from consideration.
Related Analysis
Practice All AIF-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full AIF-C01 Practice Test →