Cost-Effective CDC in S3 Data Lake
A company uses Amazon S3 to store semi-structured data in a transactional data lake. Some of the data files are small, but other data files are tens of terabytes. A data engineer must perform a change data capture (CDC) operation to identify changed data from the data source. The data source sends a full snapshot as a JSON file every day and ingests the changed data into the data lake. Which solution will capture the changed data MOST cost-effectively?
Community Votes
100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The core concept tested is managing semi-structured data lifecycle in S3; the common trap is assuming AWS-managed services like Lambda or RDS are always the default answer, ignoring the scalability and cost benefits of open-source formats for large-scale CDC.
This question addresses the most cost-effective method for performing Change Data Capture (CDC) on a transactional data lake stored in Amazon S3, specifically handling both small and large files. The correct approach utilizes open-source table formats to manage schema evolution and updates directly within the data lake.
Many candidates choose Option A (AWS Lambda), believing that serverless compute is the standard AWS way to process data. However, writing custom logic to merge terabytes of JSON data in Lambda is inefficient, costly due to execution time, and difficult to maintain compared to declarative open-source formats.
Community Discussion (7 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option C is the correct answer because open-source data lake formats (such as Apache Iceberg, Delta Lake, or Hudi) are specifically designed to handle ACID transactions, schema evolution, and upserts on data stored in S3. These formats allow you to efficiently merge incoming changes with existing data without needing to move data into a relational database first. This approach is highly scalable and cost-effective for handling both small and tens-of-terabyte files, as it leverages S3's storage efficiency and avoids the overhead of managing additional infrastructure.Why the Other Options Are Wrong
Option A (Lambda) is not cost-effective for processing tens of terabytes of data; Lambda has timeout limits and high costs for long-running, heavy data transformations. Option B (RDS) and Option D (Aurora Serverless) involve moving data out of S3 into a managed SQL database, which adds significant unnecessary complexity and cost (storage, compute, and DMS licensing/usage) when the goal is to update a data lake, not a transactional database. DMS is typically used for migrating to a database or from a database, not for updating an S3-based data lake using CDC in this context.Community Comment Notes
The community overwhelmingly supports Option C, recognizing that while AWS services are powerful, open-source table formats are the industry standard for modern data lakes. As noted by user GiorgioGss, implementing CDC-based upserts using formats like Apache Iceberg with AWS Glue is a well-documented best practice. Another user highlighted that while Lambda seems like a simple AWS service choice, the engineering effort and cost for large datasets make open-source APIs far superior. Some users initially considered AWS-centric options but were convinced by the specific mention of "semi-structured" and "transactional" requirements that favor table formats.Official Reference
Exam Strategy
When a question specifies 'data lake' and 'semi-structured data' with large volumes, look for solutions involving S3-native technologies like Glue, Athena, or open-source table formats. Avoid forcing data into RDS/Aurora unless the question explicitly requires strong relational constraints or low-latency OLTP queries.
Frequently Asked Questions
Why not use AWS Lambda for CDC?
Lambda is not cost-effective or scalable for processing tens of terabytes of JSON data due to execution time limits and high compute costs.
What are open source data lake formats?
Formats like Apache Iceberg, Delta Lake, or Hudi provide ACID transactions, schema evolution, and efficient upserts on S3 data.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →