AWS Glue Data Catalog Incremental Updates with SQS

Answer Correct answer: A, C — Use an S3 event-based AWS Glue crawler consuming SQS events and an AWS Lambda function to directly update the Data Catalog based on received events.

A data engineer configured an AWS Glue Data Catalog for data that is stored in Amazon S3 buckets. The data engineer needs to configure the Data Catalog to receive incremental updates. The data engineer sets up event notifications for the S3 bucket and creates an Amazon Simple Queue Service (Amazon SQS) queue to receive the S3 events. Which combination of steps should the data engineer take to meet these requirements with LEAST operational overhead? (Choose two.)

  1. Create an S3 event-based AWS Glue crawler to consume events from the SQS queue. Correct Answer
  2. Define a time-based schedule to run the AWS Glue crawler, and perform incremental updates to the Data Catalog.
  3. Use an AWS Lambda function to directly update the Data Catalog based on S3 events that the SQS queue receives. Correct Answer
  4. Manually initiate the AWS Glue crawler to perform updates to the Data Catalog when there is a change in the S3 bucket.
  5. Use AWS Step Functions to orchestrate the process of updating the Data Catalog based on S3 events that the SQS queue receives.

Community Votes

AC
59%
AB
41%

59% of anonymous learners picked answer AC. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The exam tests the ability to select low-maintenance, automated solutions for data cataloging, specifically distinguishing between scheduled, manual, and event-driven approaches.

This question addresses configuring AWS Glue Data Catalog for incremental updates using S3 event notifications and SQS. The solution leverages event-driven architecture to minimize operational overhead.

Many candidates choose Option B (Time-based schedule) because they believe periodic crawling is necessary for completeness, failing to recognize that event-driven crawlers or Lambda functions offer lower operational overhead for real-time needs.

Community Discussion (7 comments)

Ell89 👍 1 Selected: AC
• A leverages the event-driven capability of Glue Crawlers. • C uses AWS Lambda for direct and real-time updates to the Data Catalog. • This combination ensures incremental updates are made only when changes occur, reducing costs and operational complexity.
YUICH 👍 1 Selected: AB
(A) S3 Event-Based Crawler: Automatically triggers incremental catalog updates whenever new data arrives in the S3 bucket, reducing the need for custom code and manual intervention. (B) Time-Based Schedule: Periodically runs the crawler to catch any missed events and keep the data catalog accurate and up to date. Using both methods minimizes operational overhead while ensuring comprehensive and reliable incremental updates.
axantroff 👍 1 Selected: AB
Check out the design pattern documentation for this case. There's no need for Lambda here, so option C should be excluded. Option B seems viable, along with option A (A is the obvious choice for me). https://aws.amazon.com/blogs/big-data/run-aws-glue-crawlers-using-amazon-s3-event-notifications/
michele_scar 👍 3 Selected: AC
B and D are wrong due too "Manually" and "Scheduling". E is too much for this use case
tucobbad 👍 3 Selected: AC
  • Option A suggests creating an S3 event-based AWS Glue crawler to consume events from the SQS queue. This option is appropriate as it allows the crawler to automatically respond to events, thereby reducing manual intervention and ensuring timely updates to the Data Catalog - Option C involves using an AWS Lambda function to directly update the Data Catalog based on S3 events received from the SQS queue. This is a strong candidate as it automates the update process without the need for manual scheduling or intervention, thus minimizing operational overhead. AWS Glue Crawlers can consume events from an SQS queue: https://docs.aws.amazon.com/glue/latest/dg/crawler-s3-event-notifications.html
pikuantne 👍 3 Selected: AB
Based on this article (Option 1 for the architecture) it should be AB: 1. Run the crawler on a schedule. 2. Crawler polls for object create events in the SQS queue 3a. If there are events, crawler updates the Data Catalog 3b. If not, crawler stops
ae35a02 👍 1 Selected: BC
AWS Glue Crawlers can not consupe events from an SQS queue D introduce a manual operation E introduce more complexity so BC

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option A is correct because AWS Glue Crawlers can be configured to consume events from an Amazon SQS queue. This allows the crawler to run automatically only when new data arrives in S3, ensuring the Data Catalog is updated incrementally without manual intervention or fixed schedules. Option C is correct because using an AWS Lambda function triggered by SQS provides a serverless, event-driven mechanism to update the Data Catalog directly. This approach has minimal operational overhead as it requires no infrastructure management and runs only when needed. Together, these options represent the most efficient ways to handle incremental updates based on S3 events.

Why the Other Options Are Wrong

Option B involves a time-based schedule, which introduces operational overhead due to fixed intervals and potential unnecessary runs, contrary to the 'least operational overhead' requirement. Option D suggests manual initiation, which directly contradicts the goal of automation and minimal overhead. Option E uses AWS Step Functions, which adds significant complexity and orchestration overhead compared to the simpler Lambda or Glue Crawler triggers, making it less optimal for this specific use case.

Community Comment Notes

Several users like Michele Scar and Ell89 support AC, noting that B and D are wrong due to manual/scheduled aspects and E is too complex. User Tucobbad explains that A allows automatic response to events while C offers direct real-time updates. However, user Pikuantne and Axantroff argue for AB, citing AWS documentation about running crawlers on a schedule or polling SQS. User Ae35a02 incorrectly suggests BC, likely misunderstanding crawler capabilities. The consensus among high-voted comments favors the event-driven nature of A and C.

Official Reference

Exam Strategy

Focus on identifying keywords like 'least operational overhead' and 'incremental updates'. Prefer serverless, event-driven solutions (Lambda, Event-driven Crawlers) over scheduled or manual ones when real-time or near-real-time updates are implied by event notifications.

Frequently Asked Questions

Why not use a scheduled Glue Crawler (Option B)?

Scheduled crawlers run at fixed intervals, causing operational overhead and potential redundant executions. Event-driven crawlers run only when changes occur, aligning better with 'least operational overhead'.

Can AWS Glue Crawlers consume directly from SQS?

Yes, AWS Glue Crawlers can be configured to trigger based on Amazon SQS messages, allowing them to process S3 events asynchronously without polling.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide