Which AWS Glue Feature Enables Incremental ETL Ingestion?
A data engineer is building an automated extract, transform, and load (ETL) ingestion pipeline by using AWS Glue. The pipeline ingests compressed files that are in an Amazon S3 bucket. The ingestion pipeline must support incremental data processing. Which AWS Glue feature should the data engineer use to meet this requirement?
Community Votes
100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The exam tests whether you know which Glue feature persists processing state across runs; the trap is confusing orchestration features like Triggers and Workflows with incremental data tracking.
AWS Glue job bookmarks provide the state tracking that lets an ETL job process only new or modified objects in Amazon S3 on each run. This page establishes why Job bookmarks (C) — not workflows, triggers, or classifiers — satisfy the incremental data processing requirement for compressed-file ingestion.
Many candidates pick Triggers because triggers make jobs run automatically, but a trigger only schedules or fires a job — without job bookmarks the job still reprocesses the entire S3 prefix every time.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Job bookmarks are AWS Glue's built-in state-tracking mechanism: each job run persists the last processed S3 path, timestamp, or partition so the next run skips data already handled and reads only new or changed objects. That is exactly what the question demands when it says the pipeline "must support incremental data processing" for compressed files landing in Amazon S3. Glue natively supports bookmarks for S3 sources, including compressed formats such as gzip and bzip2, so no custom checkpointing code is needed. The transformation context stores the bookmark, making resume-after-failure behavior automatic. Because only new data is read, the job also becomes more cost-efficient and faster on repeat runs.Why the Other Options Are Wrong
Workflows orchestrate a graph of crawlers, jobs, and triggers but hold no per-source state, so they cannot decide what has already been ingested. Triggers cause a crawler or job to start on a schedule, on demand, or on an event — they answer "when does the job run," not "which data still needs processing," so every run would re-read the whole prefix. Classifiers are pattern definitions used by Glue crawlers to infer schemas and formats (CSV, JSON, XML, Grok), and they have nothing to do with tracking processed objects or incremental loads. None of these three options replaces the bookmark's persisted continuation state.Community Comment Notes
Every recorded vote chose C, and the explanations converge on the same reasoning. Bmaster supplied the authoritative AWS documentation link on monitoring continuations, which is the canonical source for job bookmark behavior. As andrologin put it, "C AWS GLue bookmarks are used to implement incremental processing." Ja13 expanded that bookmarks track the last processed state in Amazon S3 so a job resumes where it left off and only new or modified data since the last successful run is processed. HunkyBunky simply confirmed "C - is right." There is no dissent to reconcile here.Official Reference
Exam Strategy
When a Glue question mentions incremental processing, resuming, or avoiding reprocessing of already-ingested S3 data, look immediately for Job bookmarks as the answer. Triggers and Workflows are scheduling/orchestration distractors and Classifiers is a crawler-schema distractor, so map the requirement keyword to the right feature family before reading the options.
Frequently Asked Questions
Why do AWS Glue triggers not satisfy the incremental processing requirement?
Triggers only start a crawler or job on a schedule or event; they do not remember which S3 objects were already processed, so each run would reprocess the entire dataset without job bookmarks.
Do Glue job bookmarks work with compressed files in Amazon S3?
Yes. Bookmark state is tracked per transformation regardless of file compression, so gzip or bzip2 objects already read are skipped and only new or changed files are ingested.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →