Fixing Class Imbalance with the SageMaker Data Wrangler Balance Data Oversampling Operation
Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. Before the ML engineer trains the model, the ML engineer must resolve the issue of the imbalanced data. Which solution will meet this requirement with the LEAST operational effort?
Community Votes
100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
Data Wrangler includes a balance data operation that directly oversamples the minority class, and it also offers random undersampling and SMOTE, so the fix is a configured transform rather than code, while DataBrew has no equivalent built-in rebalancing transform.
Before training the fraud detection model, the engineer must resolve the class imbalance in the dataset with the least operational effort. Both AWS Glue DataBrew and SageMaker Data Wrangler offer low-code data preparation, so the deciding factor is which one provides a built-in transform for the balancing operation itself.
Choosing AWS Glue DataBrew because it is also a no-code data preparation tool, or choosing Athena and manually analyzing the imbalance. Both add effort, because neither offers the built-in oversampling transform that Data Wrangler provides for exactly this purpose.
Community Discussion (3 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The requirement is to resolve class imbalance before training with the least operational effort, and the relevant capability is oversampling the minority class. SageMaker Data Wrangler provides a balance data operation built for exactly this, supporting random oversampling, random undersampling, and SMOTE, so the engineer configures a transform rather than writing balancing logic. The operation is applied in the same low-code visual environment already used for the rest of the data preparation, which is where the low operational effort comes from. The vote was unanimous at 100 for D. GiorgioGss linked the AWS blog post on balancing your data with Data Wrangler, Sadrik noted the balance data operation by name, and ninomfr64 made the decisive distinction that DataBrew has no built-in rebalancing transform.Why the Other Options Are Wrong
Using AWS Glue DataBrew built-in features to oversample the minority class (C) is the closest competitor because DataBrew is also a low-code preparation tool, but it has no built-in transform for balancing a dataset, so the oversampling must be built manually and therefore costs more operational effort. Using Amazon Athena to identify patterns that contribute to the imbalance and adjusting the dataset accordingly (A) is an analysis step rather than a rebalancing mechanism; Athena can query the data but provides no oversampling operation, so the engineer still has to implement the fix. Using SageMaker Studio Classic built-in algorithms to process the imbalanced dataset (B) confuses model training with data preparation; the built-in algorithms are for fitting models, and training an algorithm on imbalanced data does not correct the imbalance in the dataset itself.Community Comment Notes
The community was unanimous at 100 for D, with full agreement across all substance comments. ninomfr64 provided the sharpest analysis, acknowledging that both DataBrew and Data Wrangler are low-code tools but pointing out that only Data Wrangler provides the built-in balancing transform with random oversampling, random undersampling, and SMOTE. GiorgioGss and Sadrik both identified the balance data operation as the specific feature, and GiorgioGss cited the AWS blog post dedicated to this exact use case.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →