Least Latency for AWS Glue Data Catalog Synchronization

Answer Correct answer: C — Use code that writes data to Amazon S3 to invoke the Boto3 AWS Glue create_partition API call.

A company needs to partition the Amazon S3 storage that the company uses for a data lake. The partitioning will use a path of the S3 object keys in the following format: s3://bucket/prefix/year=2023/month=01/day=01. A data engineer must ensure that the AWS Glue Data Catalog synchronizes with the S3 storage when the company adds new partitions to the bucket. Which solution will meet these requirements with the LEAST latency?

  1. Schedule an AWS Glue crawler to run every morning.
  2. Manually run the AWS Glue CreatePartition API twice each day.
  3. Use code that writes data to Amazon S3 to invoke the Boto3 AWS Glue create_partition API call. Correct Answer
  4. Run the MSCK REPAIR TABLE command from the AWS Glue console.

Community Votes

C
100%

100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The exam tests the distinction between event-driven synchronization and scheduled/batch processes, where immediate API calls provide zero latency compared to crawlers or manual repairs.

This question tests the method for synchronizing the AWS Glue Data Catalog with Amazon S3 partitions to achieve the least latency. The correct solution involves programmatically invoking the CreatePartition API immediately upon data write.

Many candidates select D (MSCK REPAIR TABLE) because it is a standard tool for syncing partition metadata, but they overlook that it requires an external trigger and does not inherently offer the 'least latency' of an immediate programmatic call.

Community Discussion (8 comments)

rralucard_ 👍 8 Selected: C
Use code that writes data to Amazon S3 to invoke the Boto3 AWS Glue create_partition API call. This approach ensures that the Data Catalog is updated as soon as new data is written to S3, providing the least latency in reflecting new partitions.
pypelyncar 👍 4 Selected: C
By embedding the Boto3 create_partition API call within the code that writes data to S3, you achieve near real-time synchronization. The Data Catalog is updated immediately after a new partition is created in S3.
tgv 👍 2 Selected: C
The explanation could be more precise regarding the interaction with Amazon S3 and AWS Glue. The key point is that the process should be triggered immediately when new data is added to S3. This can be achieved through event-driven architecture, which indeed makes the solution intuitive and efficient.
valuedate 👍 1 Selected: C
add partition after writing the data in s3
DevoteamAnalytix 👍 1 Selected: D
It's about "synchronizing AWS Glue Data Catalog with S3". So for me it's D - using MSCK REPAIR TABLE for existing S3 partitions (https://docs.aws.amazon.com/athena/latest/ug/msck-repair-table.html)
okechi 👍 2
The answer is D
GiorgioGss 👍 1 Selected: C
It's pure event-driven so... C
atu1789 👍 1 Selected: C
C. Least latency

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option C is the only approach that guarantees the lowest possible latency by synchronizing the catalog at the exact moment data is written. By embedding the Boto3 create_partition call in the data ingestion code, the system ensures that the Glue Data Catalog is updated immediately after the S3 object is placed, eliminating any delay associated with scheduling or polling.

Why the Other Options Are Wrong

Option A introduces significant latency due to the daily schedule and the time required for the crawler to scan the bucket. Option B relies on manual intervention, which is error-prone and lacks automation. Option D, MSCK REPAIR TABLE, scans the S3 bucket to detect new partitions; while effective for bulk updates, it still requires a trigger (like EventBridge or Lambda) and involves scanning overhead, making it slower than an immediate API call.

Community Comment Notes

Community consensus strongly favors C, with users noting that this approach provides "near real-time synchronization" and is "pure event-driven." Some dissenting opinions suggest D, arguing it is the standard way to sync, but these comments fail to address the specific constraint of "least latency" which demands immediacy rather than periodic repair.

Official Reference

Exam Strategy

When a question asks for the "LEAST latency" or "real-time" solution, always look for event-driven or immediate programmatic actions over scheduled crawlers or batch processes. Avoid solutions that involve polling or manual steps if an automated, immediate alternative exists.

Frequently Asked Questions

Why isn't MSCK REPAIR TABLE the best for least latency?

MSCK REPAIR TABLE scans the entire S3 path to find new partitions, which takes time. It also requires a separate trigger, adding more latency than an immediate API call.

Can I use CloudWatch Events to trigger the Glue Crawler?

Yes, but crawlers are heavier than direct API calls. For 'least latency,' directly calling create_partition is faster than starting a full crawler job.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide