AWS Glue Data Catalog Incremental Updates with SQS
A data engineer configured an AWS Glue Data Catalog for data that is stored in Amazon S3 buckets. The data engineer needs to configure the Data Catalog to receive incremental updates. The data engineer sets up event notifications for the S3 bucket and creates an Amazon Simple Queue Service (Amazon SQS) queue to receive the S3 events. Which combination of steps should the data engineer take to meet these requirements with LEAST operational overhead? (Choose two.)
Community Votes
59% of anonymous learners picked answer AC. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The exam tests the ability to select low-maintenance, automated solutions for data cataloging, specifically distinguishing between scheduled, manual, and event-driven approaches.
This question addresses configuring AWS Glue Data Catalog for incremental updates using S3 event notifications and SQS. The solution leverages event-driven architecture to minimize operational overhead.
Many candidates choose Option B (Time-based schedule) because they believe periodic crawling is necessary for completeness, failing to recognize that event-driven crawlers or Lambda functions offer lower operational overhead for real-time needs.
Community Discussion (7 comments)
- Option A suggests creating an S3 event-based AWS Glue crawler to consume events from the SQS queue. This option is appropriate as it allows the crawler to automatically respond to events, thereby reducing manual intervention and ensuring timely updates to the Data Catalog - Option C involves using an AWS Lambda function to directly update the Data Catalog based on S3 events received from the SQS queue. This is a strong candidate as it automates the update process without the need for manual scheduling or intervention, thus minimizing operational overhead. AWS Glue Crawlers can consume events from an SQS queue: https://docs.aws.amazon.com/glue/latest/dg/crawler-s3-event-notifications.html
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option A is correct because AWS Glue Crawlers can be configured to consume events from an Amazon SQS queue. This allows the crawler to run automatically only when new data arrives in S3, ensuring the Data Catalog is updated incrementally without manual intervention or fixed schedules. Option C is correct because using an AWS Lambda function triggered by SQS provides a serverless, event-driven mechanism to update the Data Catalog directly. This approach has minimal operational overhead as it requires no infrastructure management and runs only when needed. Together, these options represent the most efficient ways to handle incremental updates based on S3 events.Why the Other Options Are Wrong
Option B involves a time-based schedule, which introduces operational overhead due to fixed intervals and potential unnecessary runs, contrary to the 'least operational overhead' requirement. Option D suggests manual initiation, which directly contradicts the goal of automation and minimal overhead. Option E uses AWS Step Functions, which adds significant complexity and orchestration overhead compared to the simpler Lambda or Glue Crawler triggers, making it less optimal for this specific use case.Community Comment Notes
Several users like Michele Scar and Ell89 support AC, noting that B and D are wrong due to manual/scheduled aspects and E is too complex. User Tucobbad explains that A allows automatic response to events while C offers direct real-time updates. However, user Pikuantne and Axantroff argue for AB, citing AWS documentation about running crawlers on a schedule or polling SQS. User Ae35a02 incorrectly suggests BC, likely misunderstanding crawler capabilities. The consensus among high-voted comments favors the event-driven nature of A and C.Official Reference
Exam Strategy
Focus on identifying keywords like 'least operational overhead' and 'incremental updates'. Prefer serverless, event-driven solutions (Lambda, Event-driven Crawlers) over scheduled or manual ones when real-time or near-real-time updates are implied by event notifications.
Frequently Asked Questions
Why not use a scheduled Glue Crawler (Option B)?
Scheduled crawlers run at fixed intervals, causing operational overhead and potential redundant executions. Event-driven crawlers run only when changes occur, aligning better with 'least operational overhead'.
Can AWS Glue Crawlers consume directly from SQS?
Yes, AWS Glue Crawlers can be configured to trigger based on Amazon SQS messages, allowing them to process S3 events asynchronously without polling.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →