AWS Glue Job Bookmarks Reprocessing Cause
A data engineer needs to debug an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The data engineer enabled the bookmark feature for the AWS Glue job. The data engineer has set the maximum concurrency for the AWS Glue job to 1. The AWS Glue job is successfully writing the output to Amazon Redshift. However, the Amazon S3 files that were loaded during previous runs of the AWS Glue job are being reprocessed by subsequent runs. What is the likely reason the AWS Glue job is reprocessing the files?
Community Votes
61% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests the mechanism of AWS Glue job bookmarks, with the common trap being confusion between IAM permissions and the script-level commit requirement for state persistence.
This page explains why an AWS Glue job reprocesses S3 files despite bookmarks being enabled, focusing on the critical role of the commit statement in tracking processed data.
Many candidates incorrectly choose Option A (missing s3:GetObjectAcl permission) because they assume a permissions issue is the primary cause of bookmark failures, whereas the absence of a commit statement is the direct script-level reason for reprocessing.
Community Discussion (13 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The correct answer is D. In AWS Glue, job bookmarks track which data has been processed to prevent duplicate loads. However, this tracking relies on the Glue context object'scommit method being explicitly called within the job script. If the script does not include job.commit, the bookmark information is never saved to the Glue metadata store, causing subsequent runs to treat all files as new.Why the Other Options Are Wrong
Option A is incorrect because while permissions are necessary for Glue to function, the specific error of reprocessing existing files is directly caused by the lack of a commit, not ACL checks. Option B is irrelevant; concurrency settings affect parallelism but do not dictate bookmark state management. Option C is incorrect; using an older Glue version does not inherently break bookmarks if the script logic is correct.Community Comment Notes
The community split between A and D, but the majority correctly identified D. As user 'proserv' noted, ensuring the job run script ends withjob.commit is required to record the timestamp and path. User 'mohamedTR' correctly pointed out that commit statements relate to transactional operations and bookmark tracking, distinguishing it from simple file permissions. Official Reference
Exam Strategy
When troubleshooting AWS Glue job bookmarks, always check the script logic first. Ensure that job.commit() is present at the end of your ETL script to persist the bookmark state.
Frequently Asked Questions
Why isn't s3:GetObjectAcl the cause?
Missing ACLs typically cause access denied errors, not silent reprocessing. Bookmark state persistence requires an explicit commit.
Does concurrency affect bookmarks?
No. Concurrency controls parallel task execution but does not influence how bookmark checkpoints are saved or retrieved.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →