Building Batch Data Ingestion from S3 with AWS Glue and Model Deployment Pipelines in SageMaker Studio
An ML engineer needs to create data ingestion pipelines and ML model deployment pipelines on AWS. All the raw data is stored in Amazon S3 buckets. Which solution will meet these requirements?
Community Votes
100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
AWS Glue is built for batch ETL over data already in S3, and SageMaker Studio is the environment for building model deployment pipelines, so the pairing matches the two halves of the requirement without introducing a streaming service the batch scenario does not need.
An ML engineer needs to create both data ingestion pipelines and ML model deployment pipelines, and all the raw data already sits in Amazon S3 buckets. The ingestion side is therefore a batch data integration problem over existing objects rather than a real-time stream ingestion problem.
Using Amazon Data Firehose for the ingestion pipelines, because Firehose is a well-known ingestion service. Firehose is for real-time streaming delivery into S3, and the raw data here is already sitting in S3 as batch objects, so it adds a streaming layer with no benefit.
Community Discussion (3 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The raw data already lives in S3, which makes this a batch data integration and ETL problem, and AWS Glue is the service designed for exactly that, with crawlers, catalog, and Spark-based jobs that read and transform existing S3 data. For the second half, SageMaker Studio is the environment in which ML practitioners build and manage model deployment pipelines. Option B therefore matches both halves of the requirement to the tool actually built for each. The vote was unanimous at 100 for B, and GiorgioGss called this the main use case for Glue, while Saransundar stated the two-part mapping explicitly: Glue for data ingestion, SageMaker Studio Classic for the model deployment pipeline.Why the Other Options Are Wrong
Using Amazon Data Firehose to create the data ingestion pipelines (A) applies a real-time streaming delivery service to batch data that is already stored, as Sadrik pointed out that Firehose is for real-time streaming ingestion while this scenario implies batch processing, so it adds a streaming layer without a source stream. Using Amazon Redshift ML to create the data ingestion pipelines (C) is a mismatch on two counts: Redshift ML runs machine learning models inside a Redshift data warehouse to generate SQL predictions, which is not a general ETL ingestion pipeline, and it presupposes the data is already loaded into Redshift rather than sitting in S3. Using Amazon Athena for ingestion and a SageMaker notebook for deployment (D) puts a query engine in the ingestion role, since Athena queries data where it already is rather than building a pipeline, and a notebook is an interactive development environment rather than a deployment pipeline tool.Community Comment Notes
The community was unanimous at 100 for B with full agreement among the comments. Sadrik supplied the decisive observation for eliminating option A, noting that Firehose is for real-time streaming ingestion while the presence of raw data already in S3 makes this a batch scenario. Saransundar and GiorgioGss both matched the two requirements to their services, with GiorgioGss noting that batch ETL over S3 is the primary Glue use case, which is the same reasoning that rules out the streaming and query-based alternatives.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →