Finding Duplicate Records in a Large Unstructured Dataset with AWS Glue FindMatches
A company has a large, unstructured dataset. The dataset includes many duplicate records across several key attributes. Which solution on AWS will detect duplicates in the dataset with the LEAST code development?
Community Votes
100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
FindMatches is a purpose-built transform that applies machine learning to find fuzzy duplicates in an existing dataset, so it needs no training data and no custom code, unlike building a model on QuickSight ML Insights.
A large unstructured dataset contains many duplicate records across several key attributes, and the goal is duplicate detection with the least code development. AWS Glue FindMatches is a managed ML transform built specifically to identify matching records in an existing dataset.
Assuming SageMaker Data Wrangler is the lowest-code option because it has a Drop duplicates transform. That transform removes exact duplicate rows, whereas this dataset requires identifying matching records across several attributes, which is FindMatches' specific job.
Community Discussion (5 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
AWS Glue FindMatches is a managed transform that uses machine learning to identify duplicate or matching records in a dataset that already exists, and it needs no labeled training data. Because it is a single managed transform inside Glue, the company adds duplicate detection with essentially no code, which is exactly what the least-code-development requirement asks for. The votes were unanimous at 80 for D, and Saransundar and feelgoodfactor both highlight that FindMatches is purpose-built for duplicates and scales to large datasets.Why the Other Options Are Wrong
Amazon SageMaker Data Wrangler (C) does offer a Drop duplicates transform, as conrad2023 noted, but that is a row-level exact-duplicate cleanup inside a broader interactive data preparation tool, not fuzzy record matching across several key attributes on a large unstructured dataset, and it pulls in the overhead of provisioning a Data Wrangler session. Amazon QuickSight ML Insights (B) is a business-intelligence visualization feature for explaining anomalies in dashboards; using it to build a custom deduplication model would mean writing that model yourself, which is the opposite of least code. Amazon Mechanical Turk (A) is a crowdsourcing service that would require paying people to compare records, which is manual, expensive, and not code-based at all.Community Comment Notes
The community was decisive here, with all 80 votes for D. GiorgioGss linked the AWS Glue launch announcement describing FindMatches as letting you identify duplicate or matching records in your dataset, and Saransundar emphasized that it works without labeled training data. The only dissenting comment, from nakidal495, expressed uncertainty rather than a substantive counterargument.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →