Cost-Effective CDC in S3 Data Lake

Answer Correct answer: C — Use an open source data lake format to merge the data source with the S3 data lake to insert the new data and update the existing data.

A company uses Amazon S3 to store semi-structured data in a transactional data lake. Some of the data files are small, but other data files are tens of terabytes. A data engineer must perform a change data capture (CDC) operation to identify changed data from the data source. The data source sends a full snapshot as a JSON file every day and ingests the changed data into the data lake. Which solution will capture the changed data MOST cost-effectively?

  1. Create an AWS Lambda function to identify the changes between the previous data and the current data. Configure the Lambda function to ingest the changes into the data lake.
  2. Ingest the data into Amazon RDS for MySQL. Use AWS Database Migration Service (AWS DMS) to write the changed data to the data lake.
  3. Use an open source data lake format to merge the data source with the S3 data lake to insert the new data and update the existing data. Correct Answer
  4. Ingest the data into an Amazon Aurora MySQL DB instance that runs Aurora Serverless. Use AWS Database Migration Service (AWS DMS) to write the changed data to the data lake.

Community Votes

C
100%

100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The core concept tested is managing semi-structured data lifecycle in S3; the common trap is assuming AWS-managed services like Lambda or RDS are always the default answer, ignoring the scalability and cost benefits of open-source formats for large-scale CDC.

This question addresses the most cost-effective method for performing Change Data Capture (CDC) on a transactional data lake stored in Amazon S3, specifically handling both small and large files. The correct approach utilizes open-source table formats to manage schema evolution and updates directly within the data lake.

Many candidates choose Option A (AWS Lambda), believing that serverless compute is the standard AWS way to process data. However, writing custom logic to merge terabytes of JSON data in Lambda is inefficient, costly due to execution time, and difficult to maintain compared to declarative open-source formats.

Community Discussion (7 comments)

GiorgioGss 👍 7 Selected: C
https://aws.amazon.com/blogs/big-data/implement-a-cdc-based-upsert-in-a-data-lake-using-apache-iceberg-and-aws-glue/
plutonash 👍 1 Selected: A
Generally, AWS questions never give preference to the others solution than an AWS service so even if C could be better the answer is A
influxy 👍 1
https://aws.amazon.com/blogs/big-data/choosing-an-open-table-format-for-your-transactional-data-lake-on-aws/
FunkyFresco 👍 2 Selected: C
Ill go with Delta or something like that. is C
certplan 👍 2
Relative to cost, here are docs for the reason for option C: https://docs.aws.amazon.com/AmazonS3/latest/dev/Welcome.html https://aws.amazon.com/blogs/big-data/ https://docs.aws.amazon.com/glue/latest/dg/welcome.html https://docs.aws.amazon.com/emr/ Here are docs for reasons the others are not correct: https://aws.amazon.com/lambda/pricing/ https://aws.amazon.com/rds/pricing/ https://aws.amazon.com/dms/pricing/
damaldon 👍 1
Answ. D You can migrate data from any MySQL-compatible database (MySQL, MariaDB, or Amazon Aurora MySQL) using AWS Database Migration Service. https://docs.aws.amazon.com/dms/latest/userguide/CHAP_Source.MySQL.html
[Removed] 👍 4 Selected: C
This is a tricky one. Although option A seems like the best choice since it uses an AWS service, I believe using Delta/Iceberg APIs would be easier than writing custom code on Lambda

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option C is the correct answer because open-source data lake formats (such as Apache Iceberg, Delta Lake, or Hudi) are specifically designed to handle ACID transactions, schema evolution, and upserts on data stored in S3. These formats allow you to efficiently merge incoming changes with existing data without needing to move data into a relational database first. This approach is highly scalable and cost-effective for handling both small and tens-of-terabyte files, as it leverages S3's storage efficiency and avoids the overhead of managing additional infrastructure.

Why the Other Options Are Wrong

Option A (Lambda) is not cost-effective for processing tens of terabytes of data; Lambda has timeout limits and high costs for long-running, heavy data transformations. Option B (RDS) and Option D (Aurora Serverless) involve moving data out of S3 into a managed SQL database, which adds significant unnecessary complexity and cost (storage, compute, and DMS licensing/usage) when the goal is to update a data lake, not a transactional database. DMS is typically used for migrating to a database or from a database, not for updating an S3-based data lake using CDC in this context.

Community Comment Notes

The community overwhelmingly supports Option C, recognizing that while AWS services are powerful, open-source table formats are the industry standard for modern data lakes. As noted by user GiorgioGss, implementing CDC-based upserts using formats like Apache Iceberg with AWS Glue is a well-documented best practice. Another user highlighted that while Lambda seems like a simple AWS service choice, the engineering effort and cost for large datasets make open-source APIs far superior. Some users initially considered AWS-centric options but were convinced by the specific mention of "semi-structured" and "transactional" requirements that favor table formats.

Official Reference

Exam Strategy

When a question specifies 'data lake' and 'semi-structured data' with large volumes, look for solutions involving S3-native technologies like Glue, Athena, or open-source table formats. Avoid forcing data into RDS/Aurora unless the question explicitly requires strong relational constraints or low-latency OLTP queries.

Frequently Asked Questions

Why not use AWS Lambda for CDC?

Lambda is not cost-effective or scalable for processing tens of terabytes of JSON data due to execution time limits and high compute costs.

What are open source data lake formats?

Formats like Apache Iceberg, Delta Lake, or Hudi provide ACID transactions, schema evolution, and efficient upserts on S3 data.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide