Which algorithm groups customers by demographics and buying patterns?

A company wants to find groups for its customers based on the customers’ demographics and buying patterns. Which algorithm should the company use to meet this requirement?

  1. K-nearest neighbors (k-NN)
  2. K-means Source Reference Answer
  3. Decision tree
  4. Support vector machine

Community Votes

B
80%
A
20%

80% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

This question tests the ability to distinguish between supervised learning (classification) and unsupervised learning (clustering), with the keyword 'find groups' being the critical indicator of a clustering task.

The question tests knowledge of unsupervised clustering algorithms, specifically K-means, for customer segmentation based on demographics and buying patterns. Community consensus strongly supports K-means as the correct choice for grouping unlabeled data into natural clusters.

Many candidates incorrectly choose K-nearest neighbors (k-NN) because it also contains the letter 'K' and involves grouping concepts, but k-NN is a supervised classification algorithm that requires labeled training data, not an unsupervised clustering method.

Community Discussion (6 comments)

chdaphne 👍 1 Selected: B
K-means is a clustering algorithm widely used for customer segmentation. It groups customers based on similarities in their demographics and buying patterns, creating distinct clusters that can be analyzed for targeted marketing strategies or personalized product offerings. This algorithm is efficient, interpretable, and works well with large datasets, making it suitable for e-commerce applications.
kopper2019 👍 1
A. K-nearest neighbors (k-NN) - classification B. K-means - clustering - groups, so B
kopper2019 👍 1 Selected: A
Let's break down why: Why K-means is correct: The company wants to find "groups" of customers → This indicates a clustering task K-means is specifically designed for grouping/clustering similar data points It works well with multiple features (demographics AND buying patterns) K-means can automatically discover natural groupings in customer data It's commonly used for customer segmentation in business applications Why other options are incorrect: A (K-nearest neighbors): This is for classification when you already have labeled data, not for discovering groups
Jessiii 👍 1 Selected: B
K-means is a clustering algorithm that groups data points into clusters based on their similarities. It is particularly well-suited for unsupervised learning tasks where the goal is to identify natural groupings within the data, such as segmenting customers based on demographics and buying patterns.
AzureDP900 👍 1 Selected: B
Answer: B. K-means The company should use K-means to group customers based on demographics and buying patterns. K-means is an unsupervised clustering algorithm that effectively partitions data into natural groups, making it ideal for discovering customer segments without prior labeling.
chris_spencer 👍 1 Selected: B
K-means is a clustering algorithm

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Understanding the Problem

The key phrase in this question is "find groups" for customers based on their characteristics. This immediately signals an unsupervised learning task, specifically clustering, where the goal is to discover natural groupings in data without pre-existing labels.

Why K-means is Correct

K-means is a classic clustering algorithm that partitions data into K distinct groups (clusters) based on feature similarity. In this scenario:

  • Customer demographics (age, income, location) and buying patterns (purchase frequency, average spend, product categories) serve as the features
  • K-means automatically discovers natural customer segments by minimizing the distance between data points within each cluster
  • The result is distinct customer groups that can be used for targeted marketing and personalized offerings

Why Other Options Are Wrong

A. K-nearest neighbors (k-NN): This is a supervised classification algorithm that requires labeled training data. It classifies new data points based on the majority class of their K nearest neighbors. Since the company wants to "find groups" (discover unknown segments), not classify customers into known categories, k-NN is inappropriate.

C. Decision tree: This is also a supervised learning algorithm used for classification or regression. It requires labeled data to learn decision rules. Decision trees don't discover natural groupings; they classify based on learned patterns from labeled examples.

D. Support vector machine: SVM is another supervised classification algorithm that finds optimal boundaries between predefined classes. Like decision trees, it requires labeled training data and cannot discover unknown customer segments.

Community Insights

As noted by community members, the distinction is clear: "K-means - clustering - groups" while the other algorithms are for classification tasks. The keyword "groups" is the giveaway that this is a clustering problem requiring an unsupervised approach.

Official Reference

Exam Strategy

When you see keywords like 'find groups,' 'segment,' or 'discover patterns' without mention of labels or categories, immediately think 'unsupervised learning' and 'clustering.' Eliminate all supervised algorithms (k-NN, decision trees, SVM, logistic regression) from consideration.

Related Analysis

Practice All AIF-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full AIF-C01 Practice Test →

← Back to AIF-C01 Study Guide