AWS Glue Triggers and Connections for ETL Pipelines

Answer Correct answer: A, D — Configure AWS Glue triggers to run the ETL jobs every hour and use AWS Glue connections to establish connectivity between the data sources and Amazon Redshift.

A data engineer is building a data pipeline on AWS by using AWS Glue extract, transform, and load (ETL) jobs. The data engineer needs to process data from Amazon RDS and MongoDB, perform transformations, and load the transformed data into Amazon Redshift for analytics. The data updates must occur every hour. Which combination of tasks will meet these requirements with the LEAST operational overhead? (Choose two.)

  1. Configure AWS Glue triggers to run the ETL jobs every hour. Correct Answer
  2. Use AWS Glue DataBrew to clean and prepare the data for analytics.
  3. Use AWS Lambda functions to schedule and run the ETL jobs every hour.
  4. Use AWS Glue connections to establish connectivity between the data sources and Amazon Redshift. Correct Answer
  5. Use the Redshift Data API to load transformed data into Amazon Redshift.

Community Votes

AD
100%

100% of anonymous learners picked answer AD. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests the integration of AWS Glue's built-in orchestration features; the trap is assuming external services like Lambda or DataBrew are needed for standard scheduled ETL tasks.

This page explains how to minimize operational overhead in AWS Glue ETL pipelines by using native triggers for scheduling and connections for secure data source access.

Many learners select C (Lambda) because they overthink scheduling, failing to realize that Glue Triggers are the managed, serverless way to handle this within the service itself.

Community Discussion (11 comments)

rralucard_ 👍 7 Selected: AD
AWS Glue triggers provide a simple and integrated way to schedule ETL jobs. By configuring these triggers to run hourly, the data engineer can ensure that the data processing and updates occur as required without the need for external scheduling tools or custom scripts. This approach is directly integrated with AWS Glue, reducing the complexity and operational overhead. AWS Glue supports connections to various data sources, including Amazon RDS and MongoDB. By using AWS Glue connections, the data engineer can easily configure and manage the connectivity between these data sources and Amazon Redshift. This method leverages AWS Glue’s built-in capabilities for data source integration, thus minimizing operational complexity and ensuring a seamless data flow from the sources to the destination (Amazon Redshift).
pypelyncar 👍 6 Selected: AD
A. Configure AWS Glue triggers to run the ETL jobs every hour. Reduced Code Complexity: Glue triggers eliminate the need to write custom code for scheduling ETL jobs. This simplifies the pipeline and reduces maintenance overhead. Scalability and Integration: Glue triggers work seamlessly with Glue ETL jobs, ensuring efficient scheduling and execution within the Glue ecosystem. D. Use AWS Glue connections to establish connectivity between the data sources and Amazon Redshift. Pre-Built Connectors: Glue connections offer pre-built connectors for various data sources like RDS and Redshift. This eliminates the need for manual configuration and simplifies data source access within the ETL jobs. Centralized Management: Glue connections are centrally managed within the Glue service, streamlining connection management and reducing operational overhead.
saransh_001 👍 2 Selected: AD
A. AWS Glue provides a built-in mechanism to trigger ETL jobs at scheduled intervals, such as every hour. Using Glue triggers minimizes the need for additional custom code or services, reducing operational overhead. D. AWS Glue connections simplify the process of establishing secure and reliable connections to various data sources (Amazon RDS, MongoDB) and the destination (Amazon Redshift). This approach reduces the need for manually configuring connection settings and makes the ETL pipeline easier to maintain.
San_Juan 👍 1 Selected: AC
A. because the question is saying that the jobs are build in Glue, and must run every hour. C. because you can run the jobs as Lambda functions every hour. B. discarted, because the question is saying that "DE" is using Glue, DataBrew is for cleaning data without code, but it seems that the "DE" is writing code for transforming the data. D. Discarted, because the connections are not directly related to the question, that it is saying that you should run every hour Glue jobs, and the connections doesn't seem relevant. E. Discarted, because is saying that the data source is RDS and MongoDB, not Redshift, so you cannot use the Redshift Data API for getting the data and transform it.
sachin 👍 1
AE D is not valid. as it shoyld be Use AWS Glue connections to establish connectivity between the data sources (including Amazon Redshift) and Glue Job
DevoteamAnalytix 👍 3 Selected: AD
I was not sure about A - But in AWS console => Glue => Triggers => Add Trigger I have found the Trigger type: "Schedule - Fire the trigger on a timer."
lucas_rfsb 👍 1 Selected: CD
I found this question actually confusing. In which step the transformation would be implemented itself? I can be wrong, but with Glue triggers we would only run the job, but not the transformation logic itself. In this way, I would go in C and D
milofficial 👍 3 Selected: AD
Not a clear question - B would kinda make sense - but AD seems to be more correct
GiorgioGss 👍 4 Selected: AD
A - this is obvious and D -https://docs.aws.amazon.com/glue/latest/dg/console-connections.html
TonyStark0122 👍 3
A. Configure AWS Glue triggers to run the ETL jobs every hour. D. Use AWS Glue connections to establish connectivity between the data sources and Amazon Redshift. Explanation: Option A: Configuring AWS Glue triggers allows the ETL jobs to be scheduled and run automatically every hour without the need for manual intervention. This reduces operational overhead by automating the data processing pipeline. Option D: Using AWS Glue connections simplifies connectivity between the data sources (Amazon RDS and MongoDB) and Amazon Redshift. Glue connections abstract away the details of connection configuration, making it easier to manage and maintain the data pipeline.
milofficial 👍 2 Selected: AB
Lambda triggers for Glue jobs make me dizzy

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Options A and D are the correct choices because they leverage AWS Glue's native capabilities to achieve the goal with minimal effort. Option A is correct because AWS Glue Triggers allow you to schedule jobs (e.g., hourly) without writing custom cron expressions or managing external scheduler instances. Option D is correct because AWS Glue Connections store database credentials and network configurations securely, enabling the ETL job to connect to Amazon RDS, MongoDB, and Redshift without hardcoding sensitive information or manually configuring network endpoints each time.

Why the Other Options Are Wrong

Option B is incorrect because AWS Glue DataBrew is a no-code visual data preparation tool, which does not fit the scenario of a data engineer building a programmatic ETL pipeline with transformations. Option C is incorrect because while Lambda can trigger Glue jobs, it introduces significant operational overhead by requiring separate function management, permissions, and code maintenance compared to native Glue Triggers. Option E is incorrect because the Redshift Data API is primarily for interactive SQL querying from applications, not for bulk loading transformed data from an ETL job, where the standard JDBC/ODBC connection via Glue is more appropriate.

Community Comment Notes

The community overwhelmingly supports AD, recognizing that Glue Triggers eliminate the need for custom scheduling code. As one user noted, "A - this is obvious and D" is the standard approach for connecting sources. Another commenter pointed out that Glue triggers have a "Schedule - Fire the trigger on a timer" option, confirming their native support for periodic execution. Some users initially considered Lambda but agreed that native triggers reduce complexity. One dissenting view suggested DataBrew, but this was dismissed as it doesn't handle programmatic transformations.

Official Reference

Exam Strategy

Always look for the most 'native' AWS service feature first when 'least operational overhead' is specified. If the service mentioned in the question (Glue) has a built-in feature for the requirement (Triggers for scheduling, Connections for DB access), choose it over integrating another service like Lambda.

Frequently Asked Questions

Why not use Lambda to trigger Glue jobs?

Lambda adds operational overhead by requiring separate function creation, IAM role management, and code maintenance. Glue Triggers are native and serverless.

Can Glue Connect to MongoDB directly?

Yes, AWS Glue supports connections to various data sources including MongoDB, allowing you to store connection details securely in Glue Connections.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide