AWS Glue FindMatches for Data Deduplication
A company is migrating a legacy application to an Amazon S3 based data lake. A data engineer reviewed data that is associated with the legacy application. The data engineer found that the legacy data contained some duplicate information. The data engineer must identify and remove duplicate information from the legacy application data. Which solution will meet these requirements with the LEAST operational overhead?
Community Votes
85% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests knowledge of AWS Glue transforms, specifically identifying that FindMatches is a dedicated ML tool for deduplication designed to reduce manual effort.
This question explores the most efficient method for deduplicating legacy data during an S3 migration, highlighting AWS Glue's built-in machine learning capabilities. The correct solution leverages managed services to minimize operational overhead compared to custom code.
Many learners choose Option A or C because they assume writing Python code with libraries like Pandas or dedupe is simpler and faster than configuring an ML model, ignoring the definition of 'operational overhead' in a managed service context.
Community Discussion (5 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option B is the correct answer because it utilizes AWS Glue, a fully managed ETL service, which significantly reduces operational overhead compared to managing custom infrastructure or scripts. The FindMatches transform is specifically designed for entity resolution and deduplication using machine learning, automating the complex task of identifying duplicates without requiring the user to write custom comparison logic.Why the Other Options Are Wrong
Options A and C involve writing custom ETL jobs in Python. While Pandas (Option A) is powerful, it requires managing the execution environment, dependencies, and scaling, which increases operational overhead. Option D attempts to use the 'dedupe' library within Glue, but this is not a standard or recommended approach for leveraging AWS's native managed features, and still carries more overhead than using the built-in FindMatches transform.Community Comment Notes
Community discussion highlights a common debate: some users argue that training an ML model adds initial effort, as noted by user _JP_ who stated, "I disagree with B. That option requires additional effort just to train the ML model." However, from an exam perspective, 'operational overhead' refers to long-term maintenance, scaling, and management, where managed services like Glue FindMatches are superior to custom scripts. User rralucard_ correctly identifies that leveraging a managed service minimizes manual setup and maintenance.Official Reference
Exam Strategy
When questions ask for the 'least operational overhead,' prioritize fully managed AWS services over custom code solutions. Even if custom code seems easier to write initially, managed services like Glue, Lambda, or Step Functions are preferred for their scalability and reduced maintenance burden.
Frequently Asked Questions
Why isn't Pandas drop_duplicates the best choice?
Pandas requires managing the compute environment and script maintenance, increasing operational overhead compared to the managed Glue FindMatches transform.
Does FindMatches require extensive training?
FindMatches uses ML but is designed to be a managed transform within Glue, reducing the need for custom model building and maintenance compared to other options.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →