Choosing LightGBM in SageMaker for Fraud Detection with Class Imbalance and Correlated Features
Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. The ML engineer needs to use an Amazon SageMaker built-in algorithm to train the model. Which algorithm should the ML engineer use to meet this requirement?
Community Votes
57% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
LightGBM is a SageMaker built-in algorithm that handles class imbalance through scale_pos_weight and models feature interdependencies natively through its tree structure, which is why the linear model underperforms on this dataset.
The fraud-detection dataset is labeled but class-imbalanced, with many interdependent features, and the algorithm fails to capture the underlying patterns. The question restricts the choice to an Amazon SageMaker built-in algorithm, so the answer is the gradient-boosted tree algorithm built for exactly these two problems.
Rejecting LightGBM because it is not one of the classic first-generation SageMaker built-ins and picking Linear Learner instead. Several commenters made this argument, but the question stresses the data characteristics rather than the simplest linear fit.
Community Discussion (21 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The dataset is labeled (so supervised learning applies) and has two named problems: class imbalance and interdependent features, plus a failure to capture underlying patterns. LightGBM, a gradient-boosted decision tree algorithm available as a SageMaker built-in algorithm, directly addresses both: it exposes scale_pos_weight to rebalance classes and its boosted tree structure captures non-linear feature interactions by construction. The comments from khchan123, Saransundar, Linux_master, and ninomfr64 all reason this way, with ninomfr64 comparing scale_pos_weight against the linear learner's positive_example_weight_mult and still favoring LightGBM.Why the Other Options Are Wrong
Linear Learner (B) is a genuine built-in and can weight positive examples, but a linear model cannot capture the feature interdependencies and complex patterns the scenario says the current algorithm is missing, so it is the weaker modeling choice here. K-means clustering (C) is unsupervised and produces clusters, not a fraud classifier, so it cannot learn the labeled fraudulent-versus-legitimate boundary at all. The Neural Topic Model (D) is a topic-modeling algorithm for text corpora, aimed at document topics rather than binary fraud classification, so it is inapplicable to this transaction dataset.Community Comment Notes
This was a split vote, 57 for A and 43 for B. The minority position, argued by abrarjahin and gulf1324, held that LightGBM is not strictly a built-in algorithm and only reaches SageMaker through a prebuilt container, whereas Linear Learner is unambiguously built in. bakju0 conceded the technical nuance while noting LightGBM runs through prebuilt containers that make it easy to use in SageMaker, so the dispute centers on how strictly to read the word built-in rather than on which model fits the data.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →