Serverless sensor ingestion with Firehose, Lambda, S3, and Athena

Answer Correct answer: B — Stream data through Kinesis Data Firehose, convert to Parquet with Lambda, store in S3, and query with Athena, eliminating server maintenance and downtime.

A flood monitoring agency has deployed more than 10,000 water-level monitoring sensors. Sensors send continuous data updates, and each update is less than 1 MB in size. The agency has a fleet of on-premises application servers. These servers receive updates from the sensors, convert the raw data into a human readable format, and write the results to an on-premises relational database server. Data analysts then use simple SQL queries to monitor the data. The agency wants to increase overall application availability and reduce the effort that is required to perform maintenance tasks. These maintenance tasks, which include updates and patches to the application servers, cause downtime. While an application server is down, data is lost from sensors because the remaining servers cannot handle the entire workload. The agency wants a solution that optimizes operational overhead and costs. A solutions architect recommends the use of AWS IoT Core to collect the sensor data. What else should the solutions architect recommend to meet these requirements?

  1. Send the sensor data to Amazon Kinesis Data Firehose. Use an AWS Lambda function to read the Kinesis Data Firehose data, convert it to .csv format, and insert it into an Amazon Aurora MySQL DB instance. Instruct the data analysts to query the data directly from the DB instance.
  2. Send the sensor data to Amazon Kinesis Data Firehose. Use an AWS Lambda function to read the Kinesis Data Firehose data, convert it to Apache Parquet format, and save it to an Amazon S3 bucket. Instruct the data analysts to query the data by using Amazon Athena. Correct Answer
  3. Send the sensor data to an Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) application to convert the data to .csv format and store it in an Amazon S3 bucket. Import the data into an Amazon Aurora MySQL DB instance. Instruct the data analysts to query the data directly from the DB instance.
  4. Send the sensor data to an Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) application to convert the data to Apache Parquet format and store it in an Amazon S3 bucket. Instruct the data analysts to query the data by using Amazon Athena.

Community Votes

B
81%
A
19%

81% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Replacing on-prem application servers with a fully serverless Firehose-to-S3-to-Athena pipeline eliminates patching downtime and the single-server bottleneck that caused sensor data loss.

A flood-monitoring agency with 10,000 sensors wants higher availability and no maintenance downtime. After collecting data with AWS IoT Core, streaming through Kinesis Data Firehose, converting to Parquet with Lambda, storing in S3, and querying with Athena removes server maintenance and data loss during patches.

Using Managed Service for Apache Flink (Options C/D) — it still runs a managed application that requires operation, and Firehose is the simpler managed ingestion/transformation path for this pattern.

Community Discussion (15 comments)

CMMC 👍 7 Selected: B
Kinesis Data Firehose is well-suited for ingesting and processing streaming data at scale, such as the continuous updates from the water-level monitoring sensors. It can reliably capture and deliver data to various destinations, including S3, without requiring additional application code. Storing the data in Apache Parquet format in S3 offers several benefits. Parquet is a columnar storage format optimized for analytics workloads, providing efficient compression and query performance. This format is suitable for data analysis and querying using tools like Athena. Using AWS Lambda to transform the data from Kinesis Data Firehose into Parquet format reduces the maintenance effort associated with managing traditional servers. Lambda automatically scales with the incoming workload, ensuring continuous data processing without downtime.
mifune 👍 5 Selected: B
Lamda functions integrates with Data Firehouse better than sending the data to Apache Flink and then implement a solution to transform the data into the Parquet Format to be sent to S3. From AWS Documentation: "With Amazon Managed Service for Apache Flink, you can use Java, Scala, Python, or SQL to process and analyze streaming data". So, Flink does not make any authomatic data transformation. The correct option is B.
dv1 👍 1 Selected: B
Managed service for apache flink cannot ingest streaming data directly. This means that anything flink is out. Best remaining answer is B.
JoeTromundo 👍 1 Selected: B
Option B is the most suitable solution as it leverages serverless and scalable services (Kinesis Data Firehose, Lambda, S3, and Athena) to handle data ingestion, transformation, and analysis with minimal operational overhead and optimized costs.
sammyhaj 👍 3 Selected: A
It says convert to human readable, that isn't Parquet, its CSV
nileshlg 👍 2
D seems to be the correct option as its a managed service
seetpt 👍 2 Selected: B
B for me
titi_r 👍 3 Selected: B
“B” seems to be the correct ans. Amazon Data Firehose can ingest data streams from IoT and convert them to into Parquet format using Lambda function. The destination of the stream can be S3. https://aws.amazon.com/firehose/ https://d1.awsstatic.com/pdp-how-it-works-assets/Product-Pate-Diagram-Amazon-Kinesis-Data-Firehose%402x.39ea068e48494676c0f4386535f85a966e9ac252.png
tushar321 👍 1
D seems a better fit as Apache flink is a managed services for steaming as well as transformation. makes things simpler
VerRi 👍 4 Selected: B
Both B and D are work. B - KDF&Lambda for data transformation D - KDA for real-time analysis
Wilson_S 👍 1 Selected: D
Using a managed service for data transformation optimizes operational overhead.
oayoade 👍 3 Selected: A
"human readable format", I go with CSV
Russs99 👍 3 Selected: B
Although option D call work, it introduces unnecessary complexity for the given scenario.
Dgix 👍 2 Selected: D
Answer is D.
Sathya 👍 1
Answer is D

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option B ingests sensor data via Kinesis Data Firehose, transforms it to columnar Parquet with a Lambda function, lands it in S3, and serves analysts through Athena. The pipeline is fully serverless, so there are no application servers to patch or that can fail and drop data during maintenance.

Why the Other Options Are Wrong

Option A writes to Aurora MySQL, keeping a relational database to operate and a single point of failure. Options C and D use Managed Service for Apache Flink, which runs a managed application requiring operation and adds complexity versus Firehose. dv1 notes Flink cannot ingest streaming data directly, ruling out C/D.

Community Comment Notes

CMMC (likes 7) endorses Firehose for scalable streaming ingestion. The vote is B (74) over A (17). sammyhaj argues for CSV (A) citing 'human readable,' but Parquet with Athena meets the SQL-query requirement without servers.

Official Reference

Related Analysis

Practice All SAP-C02 Questions

Access 85 questions with complete answers and detailed explanations.

View Full SAP-C02 Practice Test →

← Back to SAP-C02 Study Guide