Processing Mixed CSV JSON XLSX Parquet Files in One S3 Folder with DataBrew and Outputting Apache Parquet
A company has an Amazon S3 bucket that contains 1 ТВ of files from different sources. The S3 bucket contains the following file types in the same S3 folder: CSV, JSON, XLSX, and Apache Parquet. An ML engineer must implement a solution that uses AWS Glue DataBrew to process the data. The ML engineer also must store the final output in Amazon S3 so that AWS Glue can consume the output in the future. Which solution will meet these requirements?
Community Votes
57% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
DataBrew can ingest the mixed-format folder directly without reorganizing it, and the output should be Apache Parquet rather than a proprietary Glue format, because Glue reads standard Apache Parquet and columnar output is what makes the future Glue consumption efficient.
An S3 bucket holds 1 TB of files from different sources with CSV, JSON, XLSX, and Apache Parquet all mixed in the same folder. The engineer must process them with AWS Glue DataBrew and store the final output in S3 so AWS Glue can consume it later, which points to a columnar output format Glue is optimized for.
Choosing a proprietary AWS Glue Parquet format, which does not exist as a distinct storage format, or splitting the 1 TB of data into separate per-format folders, which adds a large and unnecessary reorganization step for no processing benefit.
Community Discussion (5 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Two requirements drive this question: process the data with AWS Glue DataBrew, and store the output so that AWS Glue can consume it later. DataBrew can work with a dataset that contains multiple file formats in a single S3 location, so there is no need to first reorganize 1 TB of data into per-format folders, which avoids a large and pointless upfront effort. On the output side, Apache Parquet is the columnar open format that AWS Glue reads efficiently, so storing the result as Parquet satisfies the future Glue consumption requirement. The vote was 57 for A, and chris_spencer pointed out that DataBrew processes various file types within the same folder, while aws_T.server weighed in that Glue performs best with Parquet and no folder split is needed.Why the Other Options Are Wrong
Both option B and option D specify AWS Glue Parquet format, which is not a real storage format; AWS Glue consumes standard Apache Parquet, so there is no separate Glue Parquet format to write. Options C and D additionally require separating the data into a different folder for each file type before processing, which means reorganizing 1 TB of data across folders and, as eesa noted, means splitting datasets that DataBrew could otherwise read together. The choice between A and C therefore reduces to whether the unnecessary reorganization is justified, and for least effort with mixed formats in one folder it is not.Community Comment Notes
This was a close vote, 57 for A and 43 for C. ryuhei and eesa argued for C, claiming DataBrew requires a consistent format per dataset and pointing to the efficiency of Parquet, though they agreed there is no need to separate files by type, which actually undermines the separating half of option C. michele_scar favored A precisely because C implies reorganizing 1 TB of data, describing it as a lot of work with little benefit. The community is split on the DataBrew mixed-format constraint while agreeing on Parquet as the output format.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →