Least-Effort Way to Count Distinct Customers from an .xls File in S3?

Analyze data by using AWS services. Transform and process data.
Answer Correct answer: D — Use AWS Glue DataBrew to create a no-code recipe that concatenates first and last names and applies COUNT_DISTINCT.

A company receives a daily file that contains customer data in .xls format. The company stores the file in Amazon S3. The daily file is approximately 2 GB in size. A data engineer concatenates the column in the file that contains customer first names and the column that contains customer last names. The data engineer needs to determine the number of distinct customers in the file. Which solution will meet this requirement with the LEAST operational effort?

  1. Create and run an Apache Spark job in an AWS Glue notebook. Configure the job to read the S3 file and calculate the number of distinct customers.
  2. Create an AWS Glue crawler to create an AWS Glue Data Catalog of the S3 file. Run SQL queries from Amazon Athena to calculate the number of distinct customers.
  3. Create and run an Apache Spark job in Amazon EMR Serverless to calculate the number of distinct customers.
  4. Use AWS Glue DataBrew to create a recipe that uses the COUNT_DISTINCT aggregate function to calculate the number of distinct customers. Correct Answer

Community Votes

D
100%

100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests selecting the lowest-operational-effort AWS service for .xls data preparation; the trap is assuming Athena or Spark can handle .xls directly without conversion or code.

For a 2 GB .xls customer file in Amazon S3, AWS Glue DataBrew provides the least operational effort to concatenate first and last names and count distinct customers. This DEA-C01 question tests choosing a no-code AWS analytics service over Spark, Athena, or EMR Serverless.

A frequent wrong choice is B (Glue crawler plus Athena), because SQL feels simple, but Athena does not natively read .xls and the crawler, catalog, and format-conversion steps add operational effort.

Community Discussion (4 comments)

rralucard_ 👍 9 Selected: D
AWS Glue DataBrew: AWS Glue DataBrew is a visual data preparation tool that allows data engineers and data analysts to clean and normalize data without writing code. Using DataBrew, a data engineer could create a recipe that includes the concatenation of the customer first and last names and then use the COUNT_DISTINCT function. This would not require complex code and could be performed through the DataBrew user interface, representing a lower operational effort.
pypelyncar 👍 2 Selected: D
DataBrew supports various transformations, including the COUNT_DISTINCT function, which is ideal for calculating the number of unique values in a column (combined first and last names in this case).
Ousseyni 👍 2 Selected: D
go in D
lucas_rfsb 👍 2 Selected: D
since it's less operational effort, I would go in D

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option D is correct because AWS Glue DataBrew is a visual, no-code data preparation service that can read an.xls file from Amazon S3, concatenate the first-name and last-name columns, and apply the COUNT_DISTINCT aggregate function in a recipe. The engineer does not need to write Spark code, provision or tune a cluster, or build a Data Catalog and convert the file into a queryable format. That directly satisfies the requirement for the least operational effort on this 2 GB daily file. DataBrew recipes can also be saved and rerun, which fits the recurring daily pattern described in the question.

Why the Other Options Are Wrong

Option A requires writing and running an Apache Spark job in an AWS Glue notebook, which adds development and tuning effort even though Glue is serverless. Option B is tempting but Amazon Athena expects formats such as CSV, JSON, Parquet, or ORC; an.xls file must first be converted, and the Glue crawler still needs a compatible classifier, so the crawler-plus-Athena path is not least effort. Option C uses Amazon EMR Serverless with Apache Spark, which again requires custom code and job configuration for the same distinct-count logic. None of these options matches DataBrew's built-in concatenation and COUNT_DISTINCT recipe steps for the 2 GB.xls source.

Community Comment Notes

Every visible vote and comment selects D. As rralucard_ explained, DataBrew is a visual data preparation tool that can concatenate the name columns and use COUNT_DISTINCT "without writing code." pypelyncar added that COUNT_DISTINCT is ideal for calculating unique values in the combined first-and-last-name column. lucas_rfsb and Ousseyni both chose D specifically because of the "less operational effort" wording, and no comment argues for Athena, a Glue notebook, or EMR Serverless.

Official Reference

Exam Strategy

When DEA-C01 asks for the LEAST operational effort, eliminate any option that requires custom code or manual cluster and job management. Then check whether the source format is natively supported by the candidate service; for .xls files, no-code DataBrew is usually the intended answer.

Frequently Asked Questions

Why is Athena not the least-effort choice for a 2 GB .xls file in S3?

Athena does not natively read .xls; you must convert the file to a supported format and catalog it, adding operational steps compared with DataBrew.

Does AWS Glue DataBrew support .xls input and COUNT_DISTINCT?

Yes. DataBrew supports Excel files stored in Amazon S3 and provides aggregate recipe functions such as COUNT_DISTINCT for distinct customer counts.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide