Building Batch Data Ingestion from S3 with AWS Glue and Model Deployment Pipelines in SageMaker Studio

Ingest and store data.
Answer Correct answer: B — AWS Glue builds batch ingestion pipelines over raw data already in S3, and SageMaker Studio builds the model deployment pipelines, one tool per half.

An ML engineer needs to create data ingestion pipelines and ML model deployment pipelines on AWS. All the raw data is stored in Amazon S3 buckets. Which solution will meet these requirements?

  1. Use Amazon Data Firehose to create the data ingestion pipelines. Use Amazon SageMaker Studio Classic to create the model deployment pipelines.
  2. Use AWS Glue to create the data ingestion pipelines. Use Amazon SageMaker Studio Classic to create the model deployment pipelines. Correct Answer
  3. Use Amazon Redshift ML to create the data ingestion pipelines. Use Amazon SageMaker Studio Classic to create the model deployment pipelines.
  4. Use Amazon Athena to create the data ingestion pipelines. Use an Amazon SageMaker notebook to create the model deployment pipelines.

Community Votes

B
100%

100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

AWS Glue is built for batch ETL over data already in S3, and SageMaker Studio is the environment for building model deployment pipelines, so the pairing matches the two halves of the requirement without introducing a streaming service the batch scenario does not need.

An ML engineer needs to create both data ingestion pipelines and ML model deployment pipelines, and all the raw data already sits in Amazon S3 buckets. The ingestion side is therefore a batch data integration problem over existing objects rather than a real-time stream ingestion problem.

Using Amazon Data Firehose for the ingestion pipelines, because Firehose is a well-known ingestion service. Firehose is for real-time streaming delivery into S3, and the raw data here is already sitting in S3 as batch objects, so it adds a streaming layer with no benefit.

Community Discussion (3 comments)

Sadrik 👍 1 Selected: B
Amazon Kinesis Data Firehose is for real-time streaming ingestion, the question implies batch processing (raw data is in S3).
Saransundar 👍 1 Selected: B
Data ingestion - Glue ; Model deployment pipeline - sagemaker studio classic
GiorgioGss 👍 1 Selected: B
This is the main use-case for Glue.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

The raw data already lives in S3, which makes this a batch data integration and ETL problem, and AWS Glue is the service designed for exactly that, with crawlers, catalog, and Spark-based jobs that read and transform existing S3 data. For the second half, SageMaker Studio is the environment in which ML practitioners build and manage model deployment pipelines. Option B therefore matches both halves of the requirement to the tool actually built for each. The vote was unanimous at 100 for B, and GiorgioGss called this the main use case for Glue, while Saransundar stated the two-part mapping explicitly: Glue for data ingestion, SageMaker Studio Classic for the model deployment pipeline.

Why the Other Options Are Wrong

Using Amazon Data Firehose to create the data ingestion pipelines (A) applies a real-time streaming delivery service to batch data that is already stored, as Sadrik pointed out that Firehose is for real-time streaming ingestion while this scenario implies batch processing, so it adds a streaming layer without a source stream. Using Amazon Redshift ML to create the data ingestion pipelines (C) is a mismatch on two counts: Redshift ML runs machine learning models inside a Redshift data warehouse to generate SQL predictions, which is not a general ETL ingestion pipeline, and it presupposes the data is already loaded into Redshift rather than sitting in S3. Using Amazon Athena for ingestion and a SageMaker notebook for deployment (D) puts a query engine in the ingestion role, since Athena queries data where it already is rather than building a pipeline, and a notebook is an interactive development environment rather than a deployment pipeline tool.

Community Comment Notes

The community was unanimous at 100 for B with full agreement among the comments. Sadrik supplied the decisive observation for eliminating option A, noting that Firehose is for real-time streaming ingestion while the presence of raw data already in S3 makes this a batch scenario. Saransundar and GiorgioGss both matched the two requirements to their services, with GiorgioGss noting that batch ETL over S3 is the primary Glue use case, which is the same reasoning that rules out the streaming and query-based alternatives.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide