Least-Effort Way to Count Distinct Customers from an .xls File in S3?
A company receives a daily file that contains customer data in .xls format. The company stores the file in Amazon S3. The daily file is approximately 2 GB in size. A data engineer concatenates the column in the file that contains customer first names and the column that contains customer last names. The data engineer needs to determine the number of distinct customers in the file. Which solution will meet this requirement with the LEAST operational effort?
Community Votes
100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests selecting the lowest-operational-effort AWS service for .xls data preparation; the trap is assuming Athena or Spark can handle .xls directly without conversion or code.
For a 2 GB .xls customer file in Amazon S3, AWS Glue DataBrew provides the least operational effort to concatenate first and last names and count distinct customers. This DEA-C01 question tests choosing a no-code AWS analytics service over Spark, Athena, or EMR Serverless.
A frequent wrong choice is B (Glue crawler plus Athena), because SQL feels simple, but Athena does not natively read .xls and the crawler, catalog, and format-conversion steps add operational effort.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option D is correct because AWS Glue DataBrew is a visual, no-code data preparation service that can read an.xls file from Amazon S3, concatenate the first-name and last-name columns, and apply the COUNT_DISTINCT aggregate function in a recipe. The engineer does not need to write Spark code, provision or tune a cluster, or build a Data Catalog and convert the file into a queryable format. That directly satisfies the requirement for the least operational effort on this 2 GB daily file. DataBrew recipes can also be saved and rerun, which fits the recurring daily pattern described in the question.Why the Other Options Are Wrong
Option A requires writing and running an Apache Spark job in an AWS Glue notebook, which adds development and tuning effort even though Glue is serverless. Option B is tempting but Amazon Athena expects formats such as CSV, JSON, Parquet, or ORC; an.xls file must first be converted, and the Glue crawler still needs a compatible classifier, so the crawler-plus-Athena path is not least effort. Option C uses Amazon EMR Serverless with Apache Spark, which again requires custom code and job configuration for the same distinct-count logic. None of these options matches DataBrew's built-in concatenation and COUNT_DISTINCT recipe steps for the 2 GB.xls source.Community Comment Notes
Every visible vote and comment selects D. As rralucard_ explained, DataBrew is a visual data preparation tool that can concatenate the name columns and use COUNT_DISTINCT "without writing code." pypelyncar added that COUNT_DISTINCT is ideal for calculating unique values in the combined first-and-last-name column. lucas_rfsb and Ousseyni both chose D specifically because of the "less operational effort" wording, and no comment argues for Athena, a Glue notebook, or EMR Serverless.Official Reference
Exam Strategy
When DEA-C01 asks for the LEAST operational effort, eliminate any option that requires custom code or manual cluster and job management. Then check whether the source format is natively supported by the candidate service; for .xls files, no-code DataBrew is usually the intended answer.
Frequently Asked Questions
Why is Athena not the least-effort choice for a 2 GB .xls file in S3?
Athena does not natively read .xls; you must convert the file to a supported format and catalog it, adding operational steps compared with DataBrew.
Does AWS Glue DataBrew support .xls input and COUNT_DISTINCT?
Yes. DataBrew supports Excel files stored in Amazon S3 and provides aggregate recipe functions such as COUNT_DISTINCT for distinct customer counts.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →