AWS Lake Formation Aggregating Fraud-Detection Training Data from S3 and On-Premises MySQL

Ingest and store data.
Answer Correct answer: D — Lake Formation aggregates and catalogs data from S3 and on-premises MySQL into a governed data lake, unlike EMR Spark, which is a compute engine for processing.

Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. Which AWS service or feature can aggregate the data from the various data sources?

  1. Amazon EMR Spark jobs
  2. Amazon Kinesis Data Streams
  3. Amazon DynamoDB
  4. AWS Lake Formation Correct Answer

Community Votes

D
71%
A
29%

71% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Lake Formation handles data aggregation and cataloging across sources, while Amazon EMR is a compute engine that runs Spark jobs you would have to write and orchestrate yourself.

The fraud-detection training set spans Amazon S3 (transaction logs, customer profiles) and tables in an on-premises MySQL database, and the question asks which service aggregates those sources. AWS Lake Formation is built to aggregate, catalog, and manage data from many sources into a governed data lake.

Choosing Amazon EMR Spark jobs because Spark can technically read both S3 and MySQL. EMR can connect to those sources, but the question asks which service performs the aggregation, and EMR is a general-purpose compute framework rather than a data aggregation and cataloging service.

Community Discussion (21 comments)

tigrex73 👍 11 Selected: A
Amazon EMR with Spark is an excellent choice for aggregating, processing, and transforming large datasets from multiple sources (e.g., Amazon S3 and on-premises MySQL database). Spark jobs can handle both structured and unstructured. While Lake Formation is great for managing data lakes, it doesn’t provide the ETL and data processing capabilities required to aggregate and transform datasets from multiple sources.
a4002bd 👍 7 Selected: D
Is it D? AWS Lake Formation ? EMR Spark jobs is more manual.
Sadrik 👍 1 Selected: A
EMR with Spark can aggregate large datasets from multiple sources, including S3 and on-premises MySQL.
chris_spencer 👍 1 Selected: A
A. Amazon EMR Spark jobs
djeong95 👍 2 Selected: D
The answer is D (Lake Formation). This is because EMR Spark does not natively support getting data from on-prem DB as its data source. You would need DataSync or something else for that. On the other hand, Lake Formation fulfills all use cases documented clearly as links shown below. https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html#lake-formation-features https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-plan-get-data-in.html
gorloff 👍 2 Selected: D
Lake Formation ain't only the storage service - it actually an umbrella over most AWS Glue Offerings. And cause it also provides fully serverless Spark, it seems to be better option than EMR. This point is quite tricky as often when someone refers to LF, they mean the governance part only.
Udyan 👍 2 Selected: D
The correct answer is D. AWS Lake Formation. Explanation: AWS Lake Formation is designed for aggregating, organizing, and securing large datasets from multiple sources (e.g., S3, on-premises databases). It simplifies the creation of a centralized data lake, enabling seamless integration and analysis of diverse data formats. This is particularly useful for tasks like fraud detection, where data comes from different sources. Amazon EMR Spark jobs (Option A) is more suited for large-scale data processing and analytics. While it can process and transform data, it requires more operational effort to configure and manage compared to AWS Lake Formation. Why AWS Lake Formation? Aggregates and organizes data from S3 and MySQL easily. Offers integrated data cataloging for better feature engineering. Reduces operational overhead compared to setting up EMR.
shabak 👍 3 Selected: D
Lake Formation is the correct answer
dbcert87 👍 5 Selected: D
AWS Lake Formation is a service designed to aggregate, catalog, and manage data from multiple data sources, including on-premises databases and Amazon S3, making it an ideal choice for this scenario. While Amazon EMR with Apache Spark is powerful for processing and analyzing large datasets, it focuses more on data processing than on data aggregation and cataloging. It doesn't inherently manage interdependencies or schema enforcement
dbcert87 👍 1 Selected: A
Amazon EMR is correct answer for aggregation
fnuuu 👍 3 Selected: D
Data Lake is used for data discovery
xukun 👍 3 Selected: D
Once you specify where your existing databases are and provide your access credentials, Lake Formation reads the data and its metadata (schema) to understand the contents of the data source. It then imports the data to your new data lake and records the metadata in a central catalog. With Lake Formation, you can import data from MySQL, PostgreSQL, SQL Server, MariaDB, and Oracle databases running in Amazon RDS or hosted in Amazon EC2. Both bulk and incremental data loading are supported. https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html
Makendran 👍 2 Selected: A
While AWS Lake Formation could potentially be used in conjunction with other services for data lake management, Amazon EMR with Spark jobs is the most direct and powerful solution for aggregating and processing data from the various sources mentioned in this scenario. It provides the necessary tools to handle the data integration, address the class imbalance, and perform the complex feature engineering that may be required for the fraud detection model.
CloudHandsOn 👍 3 Selected: D
My first choice was Lake Formation
ninomfr64 👍 5 Selected: D
Yet another poorly worded AWS certification question. Here is my reasoning, the question is about "aggregate the data from S3 and on-premise mysql" and I do intend "aggregate" as put in the same place, therefore: A. No, while EMR spark job can connect to S3 and MySQL (spark can connect to mysql database), but it is a better tool to process data and then sore them in S3 B. No, KDS it is for delivering streaming data sources to specific destinations (S3, OpenSearch ...) C. No, DynamoDB is a nosql db that is not a great fit here D. Yes, Lake Formation "combine different types of structured and unstructured data into a centralized repository" https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html and "with Lake Formation, you can import your data using workflows" and as it is based on AWS Glue it supports both S3 and mysql
AsankaIshara 👍 3 Selected: D
Question is which AWS service or feature can aggregate the data from the various data sources? So lake formation
breathingcloud 👍 1 Selected: A
I think it is A, it is more aligned with machine learning model
AbhayD 👍 1 Selected: A
Lake formation can catalog data from various sources, it doesn't provide the data processing capabilities needed for this scenario. EMR is more appropriate in this case.
TonyKean888 👍 4 Selected: D
Data Aggregation: Lake Formation is designed to create a data lake, a centralized repository that stores and manages data from various sources, including S3, relational databases (like MySQL), and other data sources. Data Transformation: It can transform and clean data, making it suitable for analysis and machine learning. This includes handling class imbalance and feature interdependencies. Data Access: It provides a unified interface to access data, simplifying the process of integrating data from different sources into the ML model. While other options like Amazon EMR Spark jobs and Amazon Kinesis Data Streams could be used for data processing and streaming, they are not the most efficient and straightforward solutions for this specific use case. Amazon DynamoDB is a NoSQL database, not designed for batch data processing and aggregation. Therefore, AWS Lake Formation is the best choice to aggregate and prepare the data for the ML model. ref:https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html
LR2023 👍 4 Selected: D
Lake formation would be a better choice over EMR as it involes complexity setting up . For data aggregation and ETL processes, especially involving multiple data sources and ensuring data quality and security, AWS Lake Formation or Amazon Glue are more specialized and suitable option
GiorgioGss 👍 1 Selected: A
I would go with EMR Spark jobs just because I think Lake Formation is not designed for feature engineering. Spark is.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

AWS Lake Formation is the service designed to aggregate, catalog, and manage data coming from multiple locations, including Amazon S3 objects and on-premises relational databases such as MySQL. It creates a single governed data lake view over those sources, so the fraud-detection training dataset becomes centrally discoverable and usable. This matches the question's explicit requirement to aggregate data from the various data sources. Several commenters, including a4002bd and dbcert87, selected Lake Formation on exactly this aggregation-and-cataloging rationale.

Why the Other Options Are Wrong

Amazon EMR Spark jobs (A) are a legitimate way to read S3 and MySQL and process them, but EMR is a compute engine: the aggregation, joins, and placement of results depend entirely on the Spark code you write and schedule, so the service itself is not providing the data aggregation layer the question targets. Amazon Kinesis Data Streams (B) is a real-time streaming ingestion service, not a batch aggregation and cataloging tool for a static training dataset. Amazon DynamoDB (C) is a NoSQL key-value database and provides no built-in facility to pull together and catalog records from S3 and an external MySQL database.

Community Comment Notes

The majority of voters (71 of 100) chose D, and the substance comments converge on Lake Formation: a4002bd noted EMR is more manual, while TonyKean888 and LR2023 described Lake Formation as the centralized repository that stores and manages data from S3 and relational sources. ninomfr64 objected to the question's wording, arguing that Spark can reach MySQL, but still concluded Lake Formation is the intended answer.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide