How Do You Achieve Exactly-Once Delivery in a Kinesis Data Streams Pipeline?

Answer Correct answer: A — Embed a unique ID in each record at the source so the Kinesis pipeline can deduplicate events that PutRecord replays after network outages.

A banking company uses an application to collect large volumes of transactional data. The company uses Amazon Kinesis Data Streams for real-time analytics. The company’s application uses the PutRecord action to send data to Kinesis Data Streams. A data engineer has observed network outages during certain times of day. The data engineer wants to configure exactly-once delivery for the entire processing pipeline. Which solution will meet this requirement?

  1. Design the application so it can remove duplicates during processing by embedding a unique ID in each record at the source. Correct Answer
  2. Update the checkpoint configuration of the Amazon Managed Service for Apache Flink (previously known as Amazon Kinesis Data Analytics) data collection application to avoid duplicate processing of events.
  3. Design the data source so events are not ingested into Kinesis Data Streams multiple times.
  4. Stop using Kinesis Data Streams. Use Amazon EMR instead. Use Apache Flink and Apache Spark Streaming in Amazon EMR.

Community Votes

A
100%

100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests whether you know Kinesis Data Streams only guarantees at-least-once delivery and that exactly-once processing must be built with idempotent, source-tagged records — the trap is assuming duplicates can be stopped at the producer or fixed by consumer checkpointing.

Kinesis Data Streams offers at-least-once delivery, so PutRecord retries during network outages can silently write duplicate transactional records into the stream. This page establishes that exactly-once processing is achieved by embedding a unique ID in each record at the source and deduplicating downstream (A), not by preventing ingestion duplicates or by Flink checkpoint tuning.

Choosing C, 'design the data source so events are not ingested multiple times,' because it sounds like the cleanest fix. A producer cannot know whether a PutRecord actually succeeded when the network times out, so it must retry and Kinesis will store both copies — duplicate ingestion is inherent to at-least-once semantics and cannot be designed away at the source.

Community Discussion (4 comments)

Ramdi1 👍 1 Selected: A
Amazon Kinesis Data Streams does not provide exactly-once delivery natively. It ensures at-least-once delivery, meaning that under certain conditions (e.g., network failures, retries), duplicate records can occur. To achieve exactly-once processing, deduplication must be handled at the application level.
PashoQ 👍 1 Selected: A
A. Design the application so it can remove duplicates during processing by embedding a unique ID in each record at the source.
Ja13 👍 2 Selected: A
A. Design the application so it can remove duplicates during processing by embedding a unique ID in each record at the source. Explanation: Exactly-Once Delivery: Ensuring exactly-once delivery is a challenge in distributed systems, especially in the presence of network outages and retries. By embedding a unique ID in each record at the source, you can track and identify duplicate records during processing. This approach allows you to implement idempotent processing, where duplicate records can be detected and discarded, ensuring that each record is processed exactly once. De-duplication Logic: Implementing de-duplication logic based on unique IDs ensures that even if the same record is ingested multiple times due to retries or network issues, it will be processed only once by the downstream applications.
bakarys 👍 3 Selected: A
A. Design the application so it can remove duplicates during processing by embedding a unique ID in each record at the source. This approach ensures that even if a record is sent more than once due to network outages or other issues, it will only be processed once because the unique ID can be used to identify and remove any duplicates. This is a common pattern for achieving exactly-once processing semantics in distributed systems. The other options do not guarantee exactly-once delivery across the entire pipeline. Option B is partially correct but it only avoids duplicate processing within the Amazon Managed Service for Apache Flink, not across the entire pipeline. Option C is not always feasible because network issues and other factors can lead to events being ingested into Kinesis Data Streams multiple times. Option D involves changing the entire technology stack, which is not necessary to achieve the desired outcome and could introduce additional complexity and cost.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Amazon Kinesis Data Streams delivers every record at least once; a PutRecord that times out during a network outage may still have been persisted, and the application's retry writes a second copy with a new sequence number. Option A accepts that reality and makes the whole pipeline idempotent by embedding a unique ID in each record at the source, so any consumer or processing stage can detect and discard the retried copy before it affects banking balances. This is exactly the pattern AWS documents for handling duplicate records in Kinesis consumers. As Ramdi1 explains, Kinesis "does not provide exactly-once delivery natively" and "deduplication must be handled at the application level." Because the requirement is exactly-once for the entire processing pipeline, the deduplication key has to travel with the record from the producer, which is precisely what A describes.

Why the Other Options Are Wrong

Option B targets the wrong layer: Apache Flink checkpointing (via Amazon Managed Service for Apache Flink) can give exactly-once state consistency inside the Flink job, but it cannot remove duplicates that are already sitting in the Kinesis stream because PutRecord was retried upstream. Option C is unachievable in practice — when a PutRecord response is lost, the producer has no way to know whether the write succeeded, so it must retry, and Kinesis happily stores both events; you cannot design a data source that guarantees a single successful ingestion. Option D throws away the real-time architecture the company depends on and still does not solve duplication: running Apache Flink and Spark Streaming on Amazon EMR reads the same duplicated stream, so the duplicates must still be deduplicated in code.

Community Comment Notes

Community sentiment is unanimous here, which is a useful sanity check on the deduplication reasoning. bakarys argues the approach "ensures that even if a record is sent more than once due to network outages or other issues, it will only be processed once" because the unique ID identifies the copy to drop. Ja13 frames it as the standard way to handle retries, noting that "Exactly-Once Delivery: Ensuring exactly-once delivery is a challenge in distributed systems." PashoQ and the wider voting record all converge on A, and Ramdi1 supplies the underlying doctrine that Kinesis is at-least-once and deduplication belongs in the application layer.

Official Reference

Exam Strategy

When a DEA-C01 question asks for exactly-once anywhere in a Kinesis Data Streams pipeline, look for an answer that adds a unique record ID at the producer plus idempotent/deduplicating processing — and immediately eliminate options about 'preventing duplicates at the source' or 'checkpoint configuration', because Kinesis is at-least-once by design.

Frequently Asked Questions

Why can't we stop duplicate records from entering Kinesis Data Streams in the first place?

A timed-out PutRecord may have succeeded, so the application must retry and Kinesis stores both copies. Duplicate ingestion is inherent to at-least-once delivery and cannot be prevented at the source.

Does Flink checkpointing in Amazon Managed Service for Apache Flink give exactly-once delivery here?

Checkpoints guarantee exactly-once state consistency inside the Flink job, but they cannot remove duplicates that PutRecord retries already wrote into the Kinesis stream before the job read them.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide