AWS Lake Formation Aggregating Fraud-Detection Training Data from S3 and On-Premises MySQL
Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. Which AWS service or feature can aggregate the data from the various data sources?
Community Votes
71% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
Lake Formation handles data aggregation and cataloging across sources, while Amazon EMR is a compute engine that runs Spark jobs you would have to write and orchestrate yourself.
The fraud-detection training set spans Amazon S3 (transaction logs, customer profiles) and tables in an on-premises MySQL database, and the question asks which service aggregates those sources. AWS Lake Formation is built to aggregate, catalog, and manage data from many sources into a governed data lake.
Choosing Amazon EMR Spark jobs because Spark can technically read both S3 and MySQL. EMR can connect to those sources, but the question asks which service performs the aggregation, and EMR is a general-purpose compute framework rather than a data aggregation and cataloging service.
Community Discussion (21 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
AWS Lake Formation is the service designed to aggregate, catalog, and manage data coming from multiple locations, including Amazon S3 objects and on-premises relational databases such as MySQL. It creates a single governed data lake view over those sources, so the fraud-detection training dataset becomes centrally discoverable and usable. This matches the question's explicit requirement to aggregate data from the various data sources. Several commenters, including a4002bd and dbcert87, selected Lake Formation on exactly this aggregation-and-cataloging rationale.Why the Other Options Are Wrong
Amazon EMR Spark jobs (A) are a legitimate way to read S3 and MySQL and process them, but EMR is a compute engine: the aggregation, joins, and placement of results depend entirely on the Spark code you write and schedule, so the service itself is not providing the data aggregation layer the question targets. Amazon Kinesis Data Streams (B) is a real-time streaming ingestion service, not a batch aggregation and cataloging tool for a static training dataset. Amazon DynamoDB (C) is a NoSQL key-value database and provides no built-in facility to pull together and catalog records from S3 and an external MySQL database.Community Comment Notes
The majority of voters (71 of 100) chose D, and the substance comments converge on Lake Formation: a4002bd noted EMR is more manual, while TonyKean888 and LR2023 described Lake Formation as the centralized repository that stores and manages data from S3 and relational sources. ninomfr64 objected to the question's wording, arguing that Spark can reach MySQL, but still concluded Lake Formation is the intended answer.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →