Finding Duplicate Records in a Large Unstructured Dataset with AWS Glue FindMatches

Answer Correct answer: D — AWS Glue FindMatches is a managed transform that uses machine learning to find duplicate or matching records in an existing dataset with no code and no labels.

A company has a large, unstructured dataset. The dataset includes many duplicate records across several key attributes. Which solution on AWS will detect duplicates in the dataset with the LEAST code development?

  1. Use Amazon Mechanical Turk jobs to detect duplicates.
  2. Use Amazon QuickSight ML Insights to build a custom deduplication model.
  3. Use Amazon SageMaker Data Wrangler to pre-process and detect duplicates.
  4. Use the AWS Glue FindMatches transform to detect duplicates. Correct Answer

Community Votes

D
100%

100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

FindMatches is a purpose-built transform that applies machine learning to find fuzzy duplicates in an existing dataset, so it needs no training data and no custom code, unlike building a model on QuickSight ML Insights.

A large unstructured dataset contains many duplicate records across several key attributes, and the goal is duplicate detection with the least code development. AWS Glue FindMatches is a managed ML transform built specifically to identify matching records in an existing dataset.

Assuming SageMaker Data Wrangler is the lowest-code option because it has a Drop duplicates transform. That transform removes exact duplicate rows, whereas this dataset requires identifying matching records across several attributes, which is FindMatches' specific job.

Community Discussion (5 comments)

Saransundar 👍 7 Selected: D
AWS Glue FindMatches is specifically designed to identify duplicate or matching records in datasets without requiring labeled training data. It uses machine learning to find fuzzy matches and allows customization to fine-tune the matching process, making it ideal for this scenario.
GiorgioGss 👍 6 Selected: D
https://aws.amazon.com/about-aws/whats-new/2021/11/aws-glue-findmatches-new-data-existing-dataset/ "allows you to identify duplicate or matching records in your dataset"
conrad2023 👍 2 Selected: C
I would argue you can should use Data Wrangler if you want to have the least code development. See https://docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler-data-insights.html#data-wrangler-data-insights-samples "You can remove duplicate samples from the dataset using the Drop duplicates transform under Manage rows."
feelgoodfactor 👍 3 Selected: D
The AWS Glue FindMatches transform is the most appropriate solution because it is specifically designed to detect duplicates, requires minimal development effort, and scales efficiently for large datasets.
nakidal495 👍 2 Selected: A
I'm not sure but I think this is the correct answer.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

AWS Glue FindMatches is a managed transform that uses machine learning to identify duplicate or matching records in a dataset that already exists, and it needs no labeled training data. Because it is a single managed transform inside Glue, the company adds duplicate detection with essentially no code, which is exactly what the least-code-development requirement asks for. The votes were unanimous at 80 for D, and Saransundar and feelgoodfactor both highlight that FindMatches is purpose-built for duplicates and scales to large datasets.

Why the Other Options Are Wrong

Amazon SageMaker Data Wrangler (C) does offer a Drop duplicates transform, as conrad2023 noted, but that is a row-level exact-duplicate cleanup inside a broader interactive data preparation tool, not fuzzy record matching across several key attributes on a large unstructured dataset, and it pulls in the overhead of provisioning a Data Wrangler session. Amazon QuickSight ML Insights (B) is a business-intelligence visualization feature for explaining anomalies in dashboards; using it to build a custom deduplication model would mean writing that model yourself, which is the opposite of least code. Amazon Mechanical Turk (A) is a crowdsourcing service that would require paying people to compare records, which is manual, expensive, and not code-based at all.

Community Comment Notes

The community was decisive here, with all 80 votes for D. GiorgioGss linked the AWS Glue launch announcement describing FindMatches as letting you identify duplicate or matching records in your dataset, and Saransundar emphasized that it works without labeled training data. The only dissenting comment, from nakidal495, expressed uncertainty rather than a substantive counterargument.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide