Which AWS Glue Feature Enables Incremental ETL Ingestion?

Answer Correct answer: C — Job bookmarks persist each transformation's last processed S3 state so the AWS Glue ETL job reads only new or modified compressed files on subsequent runs.

A data engineer is building an automated extract, transform, and load (ETL) ingestion pipeline by using AWS Glue. The pipeline ingests compressed files that are in an Amazon S3 bucket. The ingestion pipeline must support incremental data processing. Which AWS Glue feature should the data engineer use to meet this requirement?

  1. Workflows
  2. Triggers
  3. Job bookmarks Correct Answer
  4. Classifiers

Community Votes

C
100%

100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The exam tests whether you know which Glue feature persists processing state across runs; the trap is confusing orchestration features like Triggers and Workflows with incremental data tracking.

AWS Glue job bookmarks provide the state tracking that lets an ETL job process only new or modified objects in Amazon S3 on each run. This page establishes why Job bookmarks (C) — not workflows, triggers, or classifiers — satisfy the incremental data processing requirement for compressed-file ingestion.

Many candidates pick Triggers because triggers make jobs run automatically, but a trigger only schedules or fires a job — without job bookmarks the job still reprocesses the entire S3 prefix every time.

Community Discussion (4 comments)

andrologin 👍 2 Selected: C
C AWS GLue bookmarks are used to implement incremental processing
Ja13 👍 1 Selected: C
C. Job bookmarks Here's why job bookmarks are the appropriate feature: Incremental Processing: Job bookmarks in AWS Glue help track the last processed state of data in Amazon S3. They enable the ETL job to resume from where it left off in case of interruptions or subsequent runs, ensuring that only new or modified data since the last successful run is processed (incremental processing). Automated ETL: Job bookmarks work seamlessly within AWS Glue ETL jobs, allowing the job to efficiently manage the state of processed data without the need for manual intervention. Support for Compressed Files: AWS Glue natively supports reading compressed files from Amazon S3, so the ingestion pipeline can handle compressed data formats efficiently.
HunkyBunky 👍 1 Selected: C
C - is right
Bmaster 👍 4
C is correct answer.. https://docs.aws.amazon.com/glue/latest/dg/monitor-continuations.html

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Job bookmarks are AWS Glue's built-in state-tracking mechanism: each job run persists the last processed S3 path, timestamp, or partition so the next run skips data already handled and reads only new or changed objects. That is exactly what the question demands when it says the pipeline "must support incremental data processing" for compressed files landing in Amazon S3. Glue natively supports bookmarks for S3 sources, including compressed formats such as gzip and bzip2, so no custom checkpointing code is needed. The transformation context stores the bookmark, making resume-after-failure behavior automatic. Because only new data is read, the job also becomes more cost-efficient and faster on repeat runs.

Why the Other Options Are Wrong

Workflows orchestrate a graph of crawlers, jobs, and triggers but hold no per-source state, so they cannot decide what has already been ingested. Triggers cause a crawler or job to start on a schedule, on demand, or on an event — they answer "when does the job run," not "which data still needs processing," so every run would re-read the whole prefix. Classifiers are pattern definitions used by Glue crawlers to infer schemas and formats (CSV, JSON, XML, Grok), and they have nothing to do with tracking processed objects or incremental loads. None of these three options replaces the bookmark's persisted continuation state.

Community Comment Notes

Every recorded vote chose C, and the explanations converge on the same reasoning. Bmaster supplied the authoritative AWS documentation link on monitoring continuations, which is the canonical source for job bookmark behavior. As andrologin put it, "C AWS GLue bookmarks are used to implement incremental processing." Ja13 expanded that bookmarks track the last processed state in Amazon S3 so a job resumes where it left off and only new or modified data since the last successful run is processed. HunkyBunky simply confirmed "C - is right." There is no dissent to reconcile here.

Official Reference

Exam Strategy

When a Glue question mentions incremental processing, resuming, or avoiding reprocessing of already-ingested S3 data, look immediately for Job bookmarks as the answer. Triggers and Workflows are scheduling/orchestration distractors and Classifiers is a crawler-schema distractor, so map the requirement keyword to the right feature family before reading the options.

Frequently Asked Questions

Why do AWS Glue triggers not satisfy the incremental processing requirement?

Triggers only start a crawler or job on a schedule or event; they do not remember which S3 objects were already processed, so each run would reprocess the entire dataset without job bookmarks.

Do Glue job bookmarks work with compressed files in Amazon S3?

Yes. Bookmark state is tracked per transformation regardless of file compression, so gzip or bzip2 objects already read are skipped and only new or changed files are ingested.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide