Processing Mixed CSV JSON XLSX Parquet Files in One S3 Folder with DataBrew and Outputting Apache Parquet

Answer Correct answer: A — DataBrew processes the mixed CSV JSON XLSX and Parquet files in the existing S3 folder directly, and Apache Parquet output is the columnar format Glue consumes.

A company has an Amazon S3 bucket that contains 1 ТВ of files from different sources. The S3 bucket contains the following file types in the same S3 folder: CSV, JSON, XLSX, and Apache Parquet. An ML engineer must implement a solution that uses AWS Glue DataBrew to process the data. The ML engineer also must store the final output in Amazon S3 so that AWS Glue can consume the output in the future. Which solution will meet these requirements?

  1. Use DataBrew to process the existing S3 folder. Store the output in Apache Parquet format. Correct Answer
  2. Use DataBrew to process the existing S3 folder. Store the output in AWS Glue Parquet format.
  3. Separate the data into a different folder for each file type. Use DataBrew to process each folder individually. Store the output in Apache Parquet format.
  4. Separate the data into a different folder for each file type. Use DataBrew to process each folder individually. Store the output in AWS Glue Parquet format.

Community Votes

A
57%
C
43%

57% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

DataBrew can ingest the mixed-format folder directly without reorganizing it, and the output should be Apache Parquet rather than a proprietary Glue format, because Glue reads standard Apache Parquet and columnar output is what makes the future Glue consumption efficient.

An S3 bucket holds 1 TB of files from different sources with CSV, JSON, XLSX, and Apache Parquet all mixed in the same folder. The engineer must process them with AWS Glue DataBrew and store the final output in S3 so AWS Glue can consume it later, which points to a columnar output format Glue is optimized for.

Choosing a proprietary AWS Glue Parquet format, which does not exist as a distinct storage format, or splitting the 1 TB of data into separate per-format folders, which adds a large and unnecessary reorganization step for no processing benefit.

Community Discussion (5 comments)

aws_Tamilan 👍 1 Selected: A
🔑 Keyword: Process mixed file types with AWS Glue DataBrew & store for AWS Glue ✅ Correct Answer: A. Use DataBrew to process the existing S3 folder. Store the output in Apache Parquet format. Why? AWS Glue performs best with Parquet because it is optimized for analytical queries. No need to split data into separate folders—DataBrew can handle mixed file types. Why Others Are Wrong? ❌ B. "AWS Glue Parquet format" is not a valid term. Apache Parquet is the correct format. ❌ C & D. Separating files into different folders is unnecessary—DataBrew can process multiple formats in a single folder.
michele_scar 👍 1 Selected: A
C implies that you have to re-organize all files (1 TB is a lot). This means a lot of work. For me is A, less performance but without initial overhead of organization
eesa 👍 1 Selected: C
✅ Explanation: Problem Summary: The data in S3 is mixed file formats: CSV, JSON, XLSX, and Parquet — all in one folder. You need to use AWS Glue DataBrew to process the data. The processed data must be stored in S3 for AWS Glue to consume later. Key Considerations: DataBrew Input Requirements: DataBrew datasets must be in a consistent format (CSV, JSON, XLSX, or Parquet). DataBrew cannot process mixed formats in a single dataset. You must split the data by format. DataBrew Output Format: Apache Parquet is preferred for: Efficient storage Better performance with AWS Glue and other analytics tools Columnar storage benefits in querying and transformations "AWS Glue Parquet format" does not exist — this is a distractor in the answer options.
chris_spencer 👍 2 Selected: A
Should be A. C is incorrect because it involve separating the data by file type, which is unnecessary since DataBrew can process various file types within the same folder.
ryuhei 👍 2 Selected: C
AWS Glue DataBrew can process various file formats (CSV, JSON, XLSX, Parquet) Since DataBrew can handle datasets with multiple file formats, there is no need to separate files into different folders by type. Apache Parquet is an optimal format for AWS Glue Parquet is a columnar format, which is well-suited for AWS Glue and is efficient for later analysis and ML model training. "AWS Glue Parquet format" does not exist Options B and D mention "AWS Glue Parquet format," which is incorrect. Parquet is a standard data format and is not exclusive to AWS Glue. ✅ Conclusion: Option A is the best solution because it allows DataBrew to process all files in the existing S3 folder and store the output in Apache Parquet format, which is efficient and compatible with AWS Glue. 🚀

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Two requirements drive this question: process the data with AWS Glue DataBrew, and store the output so that AWS Glue can consume it later. DataBrew can work with a dataset that contains multiple file formats in a single S3 location, so there is no need to first reorganize 1 TB of data into per-format folders, which avoids a large and pointless upfront effort. On the output side, Apache Parquet is the columnar open format that AWS Glue reads efficiently, so storing the result as Parquet satisfies the future Glue consumption requirement. The vote was 57 for A, and chris_spencer pointed out that DataBrew processes various file types within the same folder, while aws_T.server weighed in that Glue performs best with Parquet and no folder split is needed.

Why the Other Options Are Wrong

Both option B and option D specify AWS Glue Parquet format, which is not a real storage format; AWS Glue consumes standard Apache Parquet, so there is no separate Glue Parquet format to write. Options C and D additionally require separating the data into a different folder for each file type before processing, which means reorganizing 1 TB of data across folders and, as eesa noted, means splitting datasets that DataBrew could otherwise read together. The choice between A and C therefore reduces to whether the unnecessary reorganization is justified, and for least effort with mixed formats in one folder it is not.

Community Comment Notes

This was a close vote, 57 for A and 43 for C. ryuhei and eesa argued for C, claiming DataBrew requires a consistent format per dataset and pointing to the efficiency of Parquet, though they agreed there is no need to separate files by type, which actually undermines the separating half of option C. michele_scar favored A precisely because C implies reorganizing 1 TB of data, describing it as a lot of work with little benefit. The community is split on the DataBrew mixed-format constraint while agreeing on Parquet as the output format.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide