How Do You Achieve Exactly-Once Delivery in a Kinesis Data Streams Pipeline?
A banking company uses an application to collect large volumes of transactional data. The company uses Amazon Kinesis Data Streams for real-time analytics. The company’s application uses the PutRecord action to send data to Kinesis Data Streams. A data engineer has observed network outages during certain times of day. The data engineer wants to configure exactly-once delivery for the entire processing pipeline. Which solution will meet this requirement?
Community Votes
100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests whether you know Kinesis Data Streams only guarantees at-least-once delivery and that exactly-once processing must be built with idempotent, source-tagged records — the trap is assuming duplicates can be stopped at the producer or fixed by consumer checkpointing.
Kinesis Data Streams offers at-least-once delivery, so PutRecord retries during network outages can silently write duplicate transactional records into the stream. This page establishes that exactly-once processing is achieved by embedding a unique ID in each record at the source and deduplicating downstream (A), not by preventing ingestion duplicates or by Flink checkpoint tuning.
Choosing C, 'design the data source so events are not ingested multiple times,' because it sounds like the cleanest fix. A producer cannot know whether a PutRecord actually succeeded when the network times out, so it must retry and Kinesis will store both copies — duplicate ingestion is inherent to at-least-once semantics and cannot be designed away at the source.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Amazon Kinesis Data Streams delivers every record at least once; a PutRecord that times out during a network outage may still have been persisted, and the application's retry writes a second copy with a new sequence number. Option A accepts that reality and makes the whole pipeline idempotent by embedding a unique ID in each record at the source, so any consumer or processing stage can detect and discard the retried copy before it affects banking balances. This is exactly the pattern AWS documents for handling duplicate records in Kinesis consumers. As Ramdi1 explains, Kinesis "does not provide exactly-once delivery natively" and "deduplication must be handled at the application level." Because the requirement is exactly-once for the entire processing pipeline, the deduplication key has to travel with the record from the producer, which is precisely what A describes.Why the Other Options Are Wrong
Option B targets the wrong layer: Apache Flink checkpointing (via Amazon Managed Service for Apache Flink) can give exactly-once state consistency inside the Flink job, but it cannot remove duplicates that are already sitting in the Kinesis stream because PutRecord was retried upstream. Option C is unachievable in practice — when a PutRecord response is lost, the producer has no way to know whether the write succeeded, so it must retry, and Kinesis happily stores both events; you cannot design a data source that guarantees a single successful ingestion. Option D throws away the real-time architecture the company depends on and still does not solve duplication: running Apache Flink and Spark Streaming on Amazon EMR reads the same duplicated stream, so the duplicates must still be deduplicated in code.Community Comment Notes
Community sentiment is unanimous here, which is a useful sanity check on the deduplication reasoning. bakarys argues the approach "ensures that even if a record is sent more than once due to network outages or other issues, it will only be processed once" because the unique ID identifies the copy to drop. Ja13 frames it as the standard way to handle retries, noting that "Exactly-Once Delivery: Ensuring exactly-once delivery is a challenge in distributed systems." PashoQ and the wider voting record all converge on A, and Ramdi1 supplies the underlying doctrine that Kinesis is at-least-once and deduplication belongs in the application layer.Official Reference
Exam Strategy
When a DEA-C01 question asks for exactly-once anywhere in a Kinesis Data Streams pipeline, look for an answer that adds a unique record ID at the producer plus idempotent/deduplicating processing — and immediately eliminate options about 'preventing duplicates at the source' or 'checkpoint configuration', because Kinesis is at-least-once by design.
Frequently Asked Questions
Why can't we stop duplicate records from entering Kinesis Data Streams in the first place?
A timed-out PutRecord may have succeeded, so the application must retry and Kinesis stores both copies. Duplicate ingestion is inherent to at-least-once delivery and cannot be prevented at the source.
Does Flink checkpointing in Amazon Managed Service for Apache Flink give exactly-once delivery here?
Checkpoints guarantee exactly-once state consistency inside the Flink job, but they cannot remove duplicates that PutRecord retries already wrote into the Kinesis stream before the job read them.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →