Integrating AWS Glue Jobs into SageMaker Pipelines with a Callback Step
A company has AWS Glue data processing jobs that are orchestrated by an AWS Glue workflow. The AWS Glue jobs can run on a schedule or can be launched manually. The company is developing pipelines in Amazon SageMaker Pipelines for ML model development. The pipelines will use the output of the AWS Glue jobs during the data processing phase of model development. An ML engineer needs to implement a solution that integrates the AWS Glue jobs with the pipelines. Which solution will meet these requirements with the LEAST operational overhead?
Community Votes
75% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
A SageMaker Pipelines callback step triggers an external process and suspends the pipeline execution until that process reports back, which is the documented pattern for reusing an existing AWS Glue workflow inside a pipeline.
AWS Glue data processing jobs already exist and are orchestrated by a Glue workflow, and the new SageMaker Pipelines must consume their output during data processing. The requirement is to integrate the existing Glue jobs into the pipeline with the least operational overhead, and the pipeline must wait for those jobs to finish.
Choosing processing steps that point at the Glue job ARNs, or reaching for Step Functions or EventBridge. A processing step re-implements the work inside SageMaker rather than invoking the existing Glue jobs, and the other two options add a second orchestrator instead of integrating into the one the company already uses.
Community Discussion (6 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
A callback step in SageMaker Pipelines is designed for exactly this scenario: it triggers an external process, such as an AWS Glue workflow, and pauses the pipeline execution until that process signals completion, so the pipeline can safely consume the Glue output during its data processing phase. Because the existing Glue jobs and workflow stay in place and the pipeline simply waits on them, this adds the least operational overhead. The vote was 75 for C, and Linux_master and GiorgioGss both linked the AWS blog post and example that use a Glue ETL job as a SageMaker Pipelines callback step for precisely this use case.Why the Other Options Are Wrong
Processing steps (B) re-implement the logic as a SageMaker processing job and do not invoke the company's existing Glue jobs, so it duplicates work and creates a second copy of the pipeline logic to maintain, which is the opposite of least overhead. AWS Step Functions (A) would introduce an entirely separate orchestrator alongside SageMaker Pipelines, adding another state machine to build and operate. Amazon EventBridge (D) can schedule or trigger both systems in a desired order, but event-driven scheduling gives no built-in way to pause the pipeline execution until the Glue jobs actually succeed, so the pipeline could read incomplete output.Community Comment Notes
The community leaned toward C, 75 to 25. Linux_master, GiorgioGss, and Saransundar cited the AWS blog post and the SageMaker examples page showing a Glue ETL job used as a callback step in an ML pipeline. Wonjin7 and a4002bd argued for B because the wording emphasizes least operational overhead, but the processing-step approach does not run the company's existing Glue jobs and therefore does not really integrate them.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →