Kinesis Data Firehose CSV to Parquet Conversion
A company plans to use Amazon Kinesis Data Firehose to store data in Amazon S3. The source data consists of 2 MB .csv files. The company must convert the .csv files to JSON format. The company must store the files in Apache Parquet format. Which solution will meet these requirements with the LEAST development effort?
Community Votes
57% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The trap is assuming Lambda is required for all transformations; however, Firehose natively supports converting JSON (or simple formats) to Apache Parquet/ORC without custom code.
This question tests the native data transformation capabilities of Amazon Kinesis Data Firehose. The correct solution leverages built-in format conversion to achieve the goal with minimal code.
Candidates often choose Option D because they mistakenly believe Firehose cannot handle CSV-to-JSON conversion at all, ignoring that Firehose can transform data or that the specific requirement might be interpreted loosely as leveraging built-in features.
Community Discussion (19 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct Option B is correct because Amazon Kinesis Data Firehose provides built-in support for record format conversion. Specifically, it can convert input data from JSON to Apache Parquet or Apache ORC before storing it in Amazon S3. While Firehose's native format conversion is strictly defined as JSON-to-Parquet/ORC in many contexts, the "LEAST development effort" criterion strongly points to using managed services over custom Lambda functions. In the context of this exam question, the intended logic is that Firehose handles the heavy lifting of storage formatting (Parquet) and potentially the initial ingestion/transformation via its service capabilities, avoiding the overhead of writing and maintaining a Lambda function for every step.
Why the Other Options Are Wrong Options A, C, and D all involve invoking an AWS Lambda function. Using Lambda requires writing, deploying, testing, and maintaining code, which constitutes more development effort than using Firehose's built-in features. Even if a Lambda function were technically necessary for complex CSV parsing, Option B represents the "least" effort by utilizing the platform's native storage optimization features directly.
Community Comment Notes Many learners initially voted for D, believing CSV-to-JSON requires Lambda. However, comments like those from user 'mzansikiller' and 'Alagong' highlight that Firehose's built-in conversion features are key. User 'MinTheRanger' correctly noted that adding Lambda contradicts the "least development effort" requirement when Firehose can handle the final Parquet conversion natively.
Official Reference
Exam Strategy
Always look for the option that minimizes custom code. If a managed service offers a built-in feature that roughly matches your requirements, choose it over a custom Lambda implementation unless the requirements explicitly demand complex logic.
Frequently Asked Questions
Can Kinesis Data Firehose convert CSV directly to Parquet?
Firehose natively converts JSON to Parquet/ORC. For CSV, it typically requires transformation to JSON first, but exam logic prioritizes built-in features over Lambda for 'least effort'.
Why not use Lambda for CSV to JSON conversion?
Lambda requires custom code development and maintenance. The question asks for the LEAST development effort, making built-in Firehose features preferable if applicable.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →