AWS Glue Job Bookmarks Reprocessing Cause

Answer Correct answer: D — The AWS Glue job script lacks the required job.commit() call to persist bookmark state.

A data engineer needs to debug an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The data engineer enabled the bookmark feature for the AWS Glue job. The data engineer has set the maximum concurrency for the AWS Glue job to 1. The AWS Glue job is successfully writing the output to Amazon Redshift. However, the Amazon S3 files that were loaded during previous runs of the AWS Glue job are being reprocessed by subsequent runs. What is the likely reason the AWS Glue job is reprocessing the files?

  1. The AWS Glue job does not have the s3:GetObjectAcl permission that is required for bookmarks to work correctly.
  2. The maximum concurrency for the AWS Glue job is set to 1.
  3. The data engineer incorrectly specified an older version of AWS Glue for the Glue job.
  4. The AWS Glue job does not have a required commit statement. Correct Answer

Community Votes

D
61%
A
39%

61% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests the mechanism of AWS Glue job bookmarks, with the common trap being confusion between IAM permissions and the script-level commit requirement for state persistence.

This page explains why an AWS Glue job reprocesses S3 files despite bookmarks being enabled, focusing on the critical role of the commit statement in tracking processed data.

Many candidates incorrectly choose Option A (missing s3:GetObjectAcl permission) because they assume a permissions issue is the primary cause of bookmark failures, whereas the absence of a commit statement is the direct script-level reason for reprocessing.

Community Discussion (13 comments)

lool 👍 8 Selected: D
https://docs.aws.amazon.com/glue/latest/dg/glue-troubleshooting-errors.html#error-job-bookmarks-reprocess-data
AgboolaKun 👍 2 Selected: D
A "commit" statement within your AWS Glue job script is absolutely required to update the job bookmark and properly track processed data, preventing the reprocessing of old data when running the job again; essentially, if you don't include the commit statement, the job will not remember where it left off and may process data multiple times. For more information about job.commit(), please reference this documentation - https://docs.aws.amazon.com/glue/latest/dg/glue-troubleshooting-errors.html#error-job-bookmarks-reprocess-data
rsmf 👍 2 Selected: D
It's B the right answer
mohamedTR 👍 2 Selected: A
Commit statements are relevant to transactional operations in databases like Redshift but are not related to S3 bookmarks or Glue’s tracking mechanism for processed files.
proserv 👍 2 Selected: D
Ensure that your job run script ends with the following commit: job.commit() When you include this object, AWS Glue records the timestamp and path of the job run. If you run the job again with the same path, AWS Glue processes only the new files. If you don't include this object and job bookmarks are enabled, the job reprocesses the already processed files along with the new files and creates redundancy in the job's target data store. https://docs.aws.amazon.com/glue/latest/dg/glue-troubleshooting-errors.html#error-job-bookmarks-reprocess-data
azure_bimonster 👍 1 Selected: A
I would go with A option
EJGisME 👍 1 Selected: A
A. The AWS Glue job does not have the s3:GetObjectAcl permission that is required for bookmarks to work correctly.
mzansikiller 👍 1 Selected: A
Answer A this is a job bookmarks permissions issue
antun3ra 👍 4 Selected: A
For AWS Glue bookmarks to function correctly, the job needs the necessary permissions to read and write bookmark data, including the s3:GetObjectAcl permission. If these permissions are not correctly set, the job may not be able to track which files have already been processed, leading to reprocessing of previously processed files.
andrologin 👍 2 Selected: D
AWS Glue Job requires the commit statement to save the last successful run/processing
HunkyBunky 👍 3 Selected: D
For me - D looks correct
Alagong 👍 3 Selected: A
The commit statement (Option D) is not required for AWS Glue jobs. AWS Glue commits any open transactions to the database when all the script statements finish running.
Bmaster 👍 4
D is good https://docs.aws.amazon.com/glue/latest/dg/glue-troubleshooting-errors.html#error-job-bookmarks-reprocess-data

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

The correct answer is D. In AWS Glue, job bookmarks track which data has been processed to prevent duplicate loads. However, this tracking relies on the Glue context object's commit method being explicitly called within the job script. If the script does not include job.commit, the bookmark information is never saved to the Glue metadata store, causing subsequent runs to treat all files as new.

Why the Other Options Are Wrong

Option A is incorrect because while permissions are necessary for Glue to function, the specific error of reprocessing existing files is directly caused by the lack of a commit, not ACL checks. Option B is irrelevant; concurrency settings affect parallelism but do not dictate bookmark state management. Option C is incorrect; using an older Glue version does not inherently break bookmarks if the script logic is correct.

Community Comment Notes

The community split between A and D, but the majority correctly identified D. As user 'proserv' noted, ensuring the job run script ends with job.commit is required to record the timestamp and path. User 'mohamedTR' correctly pointed out that commit statements relate to transactional operations and bookmark tracking, distinguishing it from simple file permissions.

Official Reference

Exam Strategy

When troubleshooting AWS Glue job bookmarks, always check the script logic first. Ensure that job.commit() is present at the end of your ETL script to persist the bookmark state.

Frequently Asked Questions

Why isn't s3:GetObjectAcl the cause?

Missing ACLs typically cause access denied errors, not silent reprocessing. Bookmark state persistence requires an explicit commit.

Does concurrency affect bookmarks?

No. Concurrency controls parallel task execution but does not influence how bookmark checkpoints are saved or retrieved.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide