Least Latency for AWS Glue Data Catalog Synchronization
A company needs to partition the Amazon S3 storage that the company uses for a data lake. The partitioning will use a path of the S3 object keys in the following format: s3://bucket/prefix/year=2023/month=01/day=01. A data engineer must ensure that the AWS Glue Data Catalog synchronizes with the S3 storage when the company adds new partitions to the bucket. Which solution will meet these requirements with the LEAST latency?
Community Votes
100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The exam tests the distinction between event-driven synchronization and scheduled/batch processes, where immediate API calls provide zero latency compared to crawlers or manual repairs.
This question tests the method for synchronizing the AWS Glue Data Catalog with Amazon S3 partitions to achieve the least latency. The correct solution involves programmatically invoking the CreatePartition API immediately upon data write.
Many candidates select D (MSCK REPAIR TABLE) because it is a standard tool for syncing partition metadata, but they overlook that it requires an external trigger and does not inherently offer the 'least latency' of an immediate programmatic call.
Community Discussion (8 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option C is the only approach that guarantees the lowest possible latency by synchronizing the catalog at the exact moment data is written. By embedding the Boto3create_partition call in the data ingestion code, the system ensures that the Glue Data Catalog is updated immediately after the S3 object is placed, eliminating any delay associated with scheduling or polling.Why the Other Options Are Wrong
Option A introduces significant latency due to the daily schedule and the time required for the crawler to scan the bucket. Option B relies on manual intervention, which is error-prone and lacks automation. Option D, MSCK REPAIR TABLE, scans the S3 bucket to detect new partitions; while effective for bulk updates, it still requires a trigger (like EventBridge or Lambda) and involves scanning overhead, making it slower than an immediate API call.Community Comment Notes
Community consensus strongly favors C, with users noting that this approach provides "near real-time synchronization" and is "pure event-driven." Some dissenting opinions suggest D, arguing it is the standard way to sync, but these comments fail to address the specific constraint of "least latency" which demands immediacy rather than periodic repair.Official Reference
Exam Strategy
When a question asks for the "LEAST latency" or "real-time" solution, always look for event-driven or immediate programmatic actions over scheduled crawlers or batch processes. Avoid solutions that involve polling or manual steps if an automated, immediate alternative exists.
Frequently Asked Questions
Why isn't MSCK REPAIR TABLE the best for least latency?
MSCK REPAIR TABLE scans the entire S3 path to find new partitions, which takes time. It also requires a separate trigger, adding more latency than an immediate API call.
Can I use CloudWatch Events to trigger the Glue Crawler?
Yes, but crawlers are heavier than direct API calls. For 'least latency,' directly calling create_partition is faster than starting a full crawler job.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →