Choosing LightGBM in SageMaker for Fraud Detection with Class Imbalance and Correlated Features

Choose a modeling approach.
Answer Correct answer: A — LightGBM's gradient-boosted trees handle class imbalance and capture feature interdependencies, matching this labeled fraud task; it is a SageMaker built-in.

Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. The ML engineer needs to use an Amazon SageMaker built-in algorithm to train the model. Which algorithm should the ML engineer use to meet this requirement?

  1. LightGBM Correct Answer
  2. Linear learner
  3. К-means clustering
  4. Neural Topic Model (NTM)

Community Votes

A
57%
B
43%

57% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

LightGBM is a SageMaker built-in algorithm that handles class imbalance through scale_pos_weight and models feature interdependencies natively through its tree structure, which is why the linear model underperforms on this dataset.

The fraud-detection dataset is labeled but class-imbalanced, with many interdependent features, and the algorithm fails to capture the underlying patterns. The question restricts the choice to an Amazon SageMaker built-in algorithm, so the answer is the gradient-boosted tree algorithm built for exactly these two problems.

Rejecting LightGBM because it is not one of the classic first-generation SageMaker built-ins and picking Linear Learner instead. Several commenters made this argument, but the question stresses the data characteristics rather than the simplest linear fit.

Community Discussion (21 comments)

Leo2023aws 👍 8 Selected: A
https://docs.aws.amazon.com/en_kr/sagemaker/latest/dg/lightgbm.html
aragon_saa 👍 7 Selected: B
Answer is B
eesa 👍 1 Selected: A
1. Clase desbalanceada: LightGBM (Light Gradient Boosted Machine) es muy eficaz para trabajar con datasets desbalanceados, gracias a su capacidad para ajustar los pesos de clases y usar técnicas como el weighted loss o boosting adaptativo. 2. Interdependencia entre variables: LightGBM puede capturar relaciones no lineales e interacciones entre variables gracias a su estructura basada en árboles de decisión. Esto lo hace más adecuado que modelos lineales como el Linear Learner, que solo captura relaciones lineales. 3. No captura de patrones complejos: La descripción indica que el algoritmo actual no está capturando patrones subyacentes complejos, lo que sugiere la necesidad de un modelo más robusto como LightGBM, capaz de modelar relaciones complejas y no lineales en los datos
Sadrik 👍 1 Selected: B
Fraud detection is a binary classification problem, and Linear Learner is designed for classification tasks. Although LightGBM can handle binary classification tasks, including fraud detection and it is actually widely used for fraud detection, it is not available as a built-in SageMaker algorithm. Linear Learner is a built-in SageMaker algorithm.
joy5135 👍 1 Selected: B
because linear learner is a built-in algorithm while lgbm is not
chris_spencer 👍 1 Selected: B
Linear learner is a built-in algorithm whereas LightGBM can be used via custom container.
doull 👍 1 Selected: B
Linear learner is a built-in algorithm where LightBM is not
shabak 👍 1 Selected: A
ChatGPT say it's A: LightGBM
Jacobog3 👍 1 Selected: A
Is supported by Sagemaker
abrarjahin 👍 2 Selected: B
Linear Learner is a built-in algorithm provided by SageMaker for supervised learning tasks like regression and classification. LightGBM is not a built-in algorithm in Amazon SageMaker. While it is a strong gradient-boosting algorithm, it would need to be implemented as a custom script in SageMaker, which increases operational overhead.
xukun 👍 1 Selected: A
https://docs.aws.amazon.com/en_kr/sagemaker/latest/dg/lightgbm.html
Makendran 👍 1 Selected: B
In an ideal scenario, for a problem with these characteristics (fraud detection, class imbalance, feature interdependencies, complex patterns), a tree-based ensemble method like XGBoost (which is a SageMaker built-in algorithm) would be more suitable. XGBoost can handle non-linear relationships, is robust to class imbalance with proper tuning, and can capture complex patterns in the data. However, given the options provided and the requirement to use a SageMaker built-in algorithm, the Linear learner is the best available choice among these options for this specific fraud detection task.
gulf1324 👍 2 Selected: B
A. Light BGM : It's suitable model, but not built-in model for SageMaker. Answer B. Linear learner : suitable model, built-in model for SageMaker. C. K-means clustering : groups similar data points, not suitable for classification problems, and it's unsupervised learning algorithm so doesn't fit in this case(fraud detection). D. Neural Topic Model: used for topic modeling and document classification, not suitable for fraud detection
khchan123 👍 3 Selected: A
Here's why LightGBM is the most suitable algorithm for this fraud detection task: Handling Class Imbalance: LightGBM is particularly effective at handling imbalanced datasets, which is a key issue mentioned in the problem statement. It has built-in mechanisms to deal with class imbalance. Feature Interdependencies: LightGBM can capture complex feature interactions through its tree-based structure, addressing the issue of feature interdependencies mentioned in the problem. Capturing Underlying Patterns: As an advanced gradient boosting framework, LightGBM is excellent at capturing complex patterns in data, which the current algorithm is struggling with. Suitable for Fraud Detection: LightGBM is widely used in fraud detection tasks due to its high performance and ability to handle large datasets efficiently. Handling Various Data Types: It can work well with the mix of data types likely present in transaction logs, customer profiles, and database tables.
ninomfr64 👍 2 Selected: A
We have an unbalanced dataset, this means we have labelled dataset thus we are going to use a supervised model training. This reduce options to A and B (K-means and NTM are unsupervised). Both LightGBM and Linear Learner provides hyperparameter to manage unbalanced datasets, respectively "scale-pos_weight" and "positive_example_weight_mult". I would go for LightGBM as this algorithm is more suited to handle complex relationship among features, while Linear Learner learns a linear function, or, for classification problems, a linear threshold function, and maps a vector x to an approximation of the label y.
Ell89 👍 1 Selected: B
Linear Learner. LightGBM is NOT a built in algorithm which the question asks for.
michaelcloud 👍 1 Selected: A
This is a binary classification problem so LightGBM so be used. Other algorithms are not for binary classification.
bakju0 👍 2 Selected: B
Is LightGBM Built-in? Technically, no, LightGBM is not a “built-in algorithm” in the same category as SageMaker’s core algorithms (like Linear Learner or XGBoost). However, it is supported through prebuilt containers, which makes it easy to use in SageMaker.
Saransundar 👍 3 Selected: A
A. LightGBM: Handles class imbalance; captures feature interdependencies; models complex patterns. B. Linear Learner: Limited with interdependent features; struggles with complex patterns; suitable for linear relationships. C. K-means Clustering: Unsupervised algorithm; not suitable for classification; can't handle class imbalance. D. Neural Topic Model (NTM): Designed for topic modeling; unsuitable for fraud detection; doesn't address class imbalance.
Linux_master 👍 3 Selected: A
This is a binary classification problem so LightGBM so be used. Other algorithms are not for binary classification.
GiorgioGss 👍 1 Selected: A
Light Gradient Boosting Machine is effective for handling class imbalances and feature interdependencies.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

The dataset is labeled (so supervised learning applies) and has two named problems: class imbalance and interdependent features, plus a failure to capture underlying patterns. LightGBM, a gradient-boosted decision tree algorithm available as a SageMaker built-in algorithm, directly addresses both: it exposes scale_pos_weight to rebalance classes and its boosted tree structure captures non-linear feature interactions by construction. The comments from khchan123, Saransundar, Linux_master, and ninomfr64 all reason this way, with ninomfr64 comparing scale_pos_weight against the linear learner's positive_example_weight_mult and still favoring LightGBM.

Why the Other Options Are Wrong

Linear Learner (B) is a genuine built-in and can weight positive examples, but a linear model cannot capture the feature interdependencies and complex patterns the scenario says the current algorithm is missing, so it is the weaker modeling choice here. K-means clustering (C) is unsupervised and produces clusters, not a fraud classifier, so it cannot learn the labeled fraudulent-versus-legitimate boundary at all. The Neural Topic Model (D) is a topic-modeling algorithm for text corpora, aimed at document topics rather than binary fraud classification, so it is inapplicable to this transaction dataset.

Community Comment Notes

This was a split vote, 57 for A and 43 for B. The minority position, argued by abrarjahin and gulf1324, held that LightGBM is not strictly a built-in algorithm and only reaches SageMaker through a prebuilt container, whereas Linear Learner is unambiguously built in. bakju0 conceded the technical nuance while noting LightGBM runs through prebuilt containers that make it easy to use in SageMaker, so the dispute centers on how strictly to read the word built-in rather than on which model fits the data.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide