Detecting Anomalies and Visualizing Data Quality Insights with SageMaker Data Wrangler
Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. After the data is aggregated, the ML engineer must implement a solution to automatically detect anomalies in the data and to visualize the result. Which solution will meet these requirements?
Community Votes
100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
SageMaker Data Wrangler includes a data quality and insights report that automatically surfaces anomalies, and its built-in analyses generate visualizations in a few clicks, so detection and visualization are both available in the same service without a separate BI tool.
In the fraud detection case study, after the data from S3 and on-premises MySQL has been aggregated, the engineer must automatically detect anomalies in the data and visualize the result. The question is designed to catch candidates who assume a separate BI tool is required for the visualization step.
Reaching for Amazon QuickSight, Amazon Redshift Spectrum, or Amazon Athena because the question mentions visualizing the result, and assuming visualization always requires a separate business intelligence tool. Data Wrangler's built-in analyses already produce the visualizations.
Community Discussion (3 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The requirement is to automatically detect anomalies in the aggregated data and to visualize the result. SageMaker Data Wrangler includes a data quality and insights report that identifies anomalies, and its built-in analyses generate visualizations and data analyses in a few clicks, so both halves of the requirement are met inside one service. This is the trick of the question: it reads as though a separate visualization product is required, but the Data Wrangler feature set already covers it. The vote was unanimous at 100 for C. GiorgioGss framed the question as one that makes you think QuickSight is needed for visualization, and cited the Data Wrangler analyses page, while ninomfr64 pointed to the Data Quality and Insights Report and the analyses page for the visualization side.Why the Other Options Are Wrong
Amazon Athena (A) is a serverless query engine that runs standard SQL queries over data in S3; it has no automatic anomaly detection, and it is not a visualization tool, so option A fails both halves of the requirement. Amazon Redshift Spectrum with Amazon QuickSight (B) can query and visualize, but Redshift Spectrum does not automatically detect anomalies, so the detection requirement is unmet. AWS Batch with Amazon QuickSight (D) only provides a compute environment to run code the engineer would still have to write; it has no automatic anomaly detection and no built-in visualization, so it requires the most custom development of all the options.Community Comment Notes
The community was unanimous at 100 for C, and the comments focus on the misconception the question is testing. GiorgioGss explicitly called the question tricky for making the candidate think QuickSight is required for visualization, citing the Data Wrangler analyses documentation. ninomfr64 identified the Data Quality and Insights Report as the feature that surfaces anomalies. motk123, despite appearing to paste content about transforming categorical data, still selected C.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →