AWS Glue FindMatches for Data Deduplication

Answer Correct answer: B — Write an AWS Glue extract, transform, and load (ETL) job. Use the FindMatches machine learning (ML) transform to transform the data to perform data deduplication.

A company is migrating a legacy application to an Amazon S3 based data lake. A data engineer reviewed data that is associated with the legacy application. The data engineer found that the legacy data contained some duplicate information. The data engineer must identify and remove duplicate information from the legacy application data. Which solution will meet these requirements with the LEAST operational overhead?

  1. Write a custom extract, transform, and load (ETL) job in Python. Use the DataFrame.drop_duplicates() function by importing the Pandas library to perform data deduplication.
  2. Write an AWS Glue extract, transform, and load (ETL) job. Use the FindMatches machine learning (ML) transform to transform the data to perform data deduplication. Correct Answer
  3. Write a custom extract, transform, and load (ETL) job in Python. Import the Python dedupe library. Use the dedupe library to perform data deduplication.
  4. Write an AWS Glue extract, transform, and load (ETL) job. Import the Python dedupe library. Use the dedupe library to perform data deduplication.

Community Votes

B
85%
A
15%

85% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests knowledge of AWS Glue transforms, specifically identifying that FindMatches is a dedicated ML tool for deduplication designed to reduce manual effort.

This question explores the most efficient method for deduplicating legacy data during an S3 migration, highlighting AWS Glue's built-in machine learning capabilities. The correct solution leverages managed services to minimize operational overhead compared to custom code.

Many learners choose Option A or C because they assume writing Python code with libraries like Pandas or dedupe is simpler and faster than configuring an ML model, ignoring the definition of 'operational overhead' in a managed service context.

Community Discussion (5 comments)

rralucard_ 👍 6 Selected: B
Option B, writing an AWS Glue ETL job with the FindMatches ML transform, is likely to meet the requirements with the least operational overhead. This solution leverages a managed service (AWS Glue) and incorporates a built-in ML transform specifically designed for deduplication, thus minimizing the need for manual setup, maintenance, and machine learning expertise.
_JP_ 👍 2 Selected: A
I disagree with B. That option requires additional effort just to train the ML model with labeled data. Option A is as simple as to use the robust pandas library
V0811 👍 1 Selected: B
100 % B
GiorgioGss 👍 4 Selected: B
B. https://docs.aws.amazon.com/glue/latest/dg/machine-learning.html "Find matches Finds duplicate records in the source data. You teach this machine learning transform by labeling example datasets to indicate which rows match. The machine learning transform learns which rows should be matches the more you teach it with example labeled data."
Aesthet 👍 1
Remove duplicates from already migrated data - probably D. Remove duplicates from data before migration - A is preferable.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option B is the correct answer because it utilizes AWS Glue, a fully managed ETL service, which significantly reduces operational overhead compared to managing custom infrastructure or scripts. The FindMatches transform is specifically designed for entity resolution and deduplication using machine learning, automating the complex task of identifying duplicates without requiring the user to write custom comparison logic.

Why the Other Options Are Wrong

Options A and C involve writing custom ETL jobs in Python. While Pandas (Option A) is powerful, it requires managing the execution environment, dependencies, and scaling, which increases operational overhead. Option D attempts to use the 'dedupe' library within Glue, but this is not a standard or recommended approach for leveraging AWS's native managed features, and still carries more overhead than using the built-in FindMatches transform.

Community Comment Notes

Community discussion highlights a common debate: some users argue that training an ML model adds initial effort, as noted by user _JP_ who stated, "I disagree with B. That option requires additional effort just to train the ML model." However, from an exam perspective, 'operational overhead' refers to long-term maintenance, scaling, and management, where managed services like Glue FindMatches are superior to custom scripts. User rralucard_ correctly identifies that leveraging a managed service minimizes manual setup and maintenance.

Official Reference

Exam Strategy

When questions ask for the 'least operational overhead,' prioritize fully managed AWS services over custom code solutions. Even if custom code seems easier to write initially, managed services like Glue, Lambda, or Step Functions are preferred for their scalability and reduced maintenance burden.

Frequently Asked Questions

Why isn't Pandas drop_duplicates the best choice?

Pandas requires managing the compute environment and script maintenance, increasing operational overhead compared to the managed Glue FindMatches transform.

Does FindMatches require extensive training?

FindMatches uses ML but is designed to be a managed transform within Glue, reducing the need for custom model building and maintenance compared to other options.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide