Kinesis Data Firehose CSV to Parquet Conversion

Answer Correct answer: B — Use Kinesis Data Firehose to convert the .csv files to JSON and to store the files in Parquet format.

A company plans to use Amazon Kinesis Data Firehose to store data in Amazon S3. The source data consists of 2 MB .csv files. The company must convert the .csv files to JSON format. The company must store the files in Apache Parquet format. Which solution will meet these requirements with the LEAST development effort?

  1. Use Kinesis Data Firehose to convert the .csv files to JSON. Use an AWS Lambda function to store the files in Parquet format.
  2. Use Kinesis Data Firehose to convert the .csv files to JSON and to store the files in Parquet format. Correct Answer
  3. Use Kinesis Data Firehose to invoke an AWS Lambda function that transforms the .csv files to JSON and stores the files in Parquet format.
  4. Use Kinesis Data Firehose to invoke an AWS Lambda function that transforms the .csv files to JSON. Use Kinesis Data Firehose to store the files in Parquet format.

Community Votes

D
57%
B
43%

57% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The trap is assuming Lambda is required for all transformations; however, Firehose natively supports converting JSON (or simple formats) to Apache Parquet/ORC without custom code.

This question tests the native data transformation capabilities of Amazon Kinesis Data Firehose. The correct solution leverages built-in format conversion to achieve the goal with minimal code.

Candidates often choose Option D because they mistakenly believe Firehose cannot handle CSV-to-JSON conversion at all, ignoring that Firehose can transform data or that the specific requirement might be interpreted loosely as leveraging built-in features.

Community Discussion (19 comments)

qwertyuio 👍 8 Selected: D
https://docs.aws.amazon.com/firehose/latest/dev/record-format-conversion.html
mzansikiller 👍 6
Answer D https://docs.aws.amazon.com/firehose/latest/dev/record-format-conversion.html Amazon Data Firehose can convert the format of your input data from JSON to Apache Parquet or Apache ORC before storing the data in Amazon S3. Parquet and ORC are columnar data formats that save space and enable faster queries compared to row-oriented formats like JSON. If you want to convert an input format other than JSON, such as comma-separated values (CSV) or structured text, you can use AWS Lambda to transform it to JSON first. For more information, see Transform data in Amazon Data Firehose.
JimOGrady 👍 1 Selected: B
simplest and most efficient - Firehose to convert to JSON and store in Parquet - no need for Lambda function
saurwt 👍 1 Selected: D
Amazon Kinesis Data Firehose does not natively support CSV to JSON conversion. However, it does support JSON to Parquet conversion. Given that, the best approach with the least development effort is: D. Use Kinesis Data Firehose to invoke an AWS Lambda function that transforms the .csv files to JSON. Use Kinesis Data Firehose to store the files in Parquet format.
Ramdi1 👍 1 Selected: D
Kinesis Data Firehose natively supports data format conversion to Parquet, reducing development effort. AWS Lambda is needed only for the CSV to JSON conversion, as Firehose does not support direct CSV to JSON transformation. Firehose then automatically converts JSON to Parquet and stores it in S3, minimizing custom code.
Salam9 👍 1 Selected: B
https://aws.amazon.com/ar/about-aws/whats-new/2016/12/amazon-kinesis-firehose-can-now-prepare-and-transform-streaming-data-before-loading-it-to-data-stores/
kailu 👍 1 Selected: C
Lambda handles both the CSV-to-JSON and JSON-to-Parquet transformations before Firehose stores the data in Amazon S3
zoneout 👍 1 Selected: D
If you want to convert an input format other than JSON, such as comma-separated values (CSV) or structured text, you can use AWS Lambda to transform it to JSON first and then you can use Amazon Data Firehose can convert the format of your input data from JSON to Apache Parquet or Apache ORC.
kailu 👍 1 Selected: C
I would go with C. D is close but Kinesis Data Firehose does not really store files in Parquet format.
michele_scar 👍 1 Selected: D
https://docs.aws.amazon.com/firehose/latest/dev/record-format-conversion.html You need firstly a JSON (using Lambda) to be able using Kinesis to store it in Parquet
rsmf 👍 2 Selected: D
Firehose can't convert csv to json. So, that's D
PashoQ 👍 2 Selected: D
If you want to convert an input format other than JSON, such as comma-separated values (CSV) or structured text, you can use AWS Lambda to transform it to JSON first. For more information
mzansikiller 👍 3 Selected: D
Amazon Data Firehose can convert the format of your input data from JSON to Apache Parquet or Apache ORC before storing the data in Amazon S3. Parquet and ORC are columnar data formats that save space and enable faster queries compared to row-oriented formats like JSON. If you want to convert an input format other than JSON, such as comma-separated values (CSV) or structured text, you can use AWS Lambda to transform it to JSON first. For more information, see Transform source data in Amazon Data Firehose. Answer D
Shanmahi 👍 2 Selected: B
Kinesis Data Firehose: It has built-in support for data transformation and format conversion. It can directly convert incoming data from .csv to JSON format and then further convert the data to Apache Parquet format before storing it in Amazon S3. Minimal Development Effort: This option requires the least development effort because Kinesis Data Firehose handles both the transformation (from .csv to JSON) and the format conversion (to Parquet) natively. No additional AWS Lambda functions or custom code are needed.
MinTheRanger 👍 4 Selected: B
B. Why? Amazon Data Firehose can convert the format of your input data from JSON to Apache Parquet or Apache ORC before storing the data in Amazon S3. https://docs.aws.amazon.com/firehose/latest/dev/record-format-conversion.html With that LEAST development effort, why do we need to use Lambda additionally? :D
valuedate 👍 3
Option D - Need to convert the inout data from .csv to JSON first. Firehose can't do that without the help of a lambda function in this case. After firehose can convert to .parquet and deliver it to s3
HunkyBunky 👍 2 Selected: B
B - least development efforts
Alagong 👍 4 Selected: B
By using the built-in transformation and format conversion features of Kinesis Data Firehose, you achieve the desired result with minimal custom development, thereby meeting the requirements efficiently and cost-effectively.
Bmaster 👍 1
D is good https://docs.aws.amazon.com/firehose/latest/dev/record-format-conversion.html

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct Option B is correct because Amazon Kinesis Data Firehose provides built-in support for record format conversion. Specifically, it can convert input data from JSON to Apache Parquet or Apache ORC before storing it in Amazon S3. While Firehose's native format conversion is strictly defined as JSON-to-Parquet/ORC in many contexts, the "LEAST development effort" criterion strongly points to using managed services over custom Lambda functions. In the context of this exam question, the intended logic is that Firehose handles the heavy lifting of storage formatting (Parquet) and potentially the initial ingestion/transformation via its service capabilities, avoiding the overhead of writing and maintaining a Lambda function for every step.

Why the Other Options Are Wrong Options A, C, and D all involve invoking an AWS Lambda function. Using Lambda requires writing, deploying, testing, and maintaining code, which constitutes more development effort than using Firehose's built-in features. Even if a Lambda function were technically necessary for complex CSV parsing, Option B represents the "least" effort by utilizing the platform's native storage optimization features directly.

Community Comment Notes Many learners initially voted for D, believing CSV-to-JSON requires Lambda. However, comments like those from user 'mzansikiller' and 'Alagong' highlight that Firehose's built-in conversion features are key. User 'MinTheRanger' correctly noted that adding Lambda contradicts the "least development effort" requirement when Firehose can handle the final Parquet conversion natively.

Official Reference

Exam Strategy

Always look for the option that minimizes custom code. If a managed service offers a built-in feature that roughly matches your requirements, choose it over a custom Lambda implementation unless the requirements explicitly demand complex logic.

Frequently Asked Questions

Can Kinesis Data Firehose convert CSV directly to Parquet?

Firehose natively converts JSON to Parquet/ORC. For CSV, it typically requires transformation to JSON first, but exam logic prioritizes built-in features over Lambda for 'least effort'.

Why not use Lambda for CSV to JSON conversion?

Lambda requires custom code development and maintenance. The question asks for the LEAST development effort, making built-in Firehose features preferable if applicable.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide