Most Cost-Effective AWS Glue ETL for Small Daily CSV Files?
A data engineer needs to build an extract, transform, and load (ETL) job. The ETL job will process daily incoming .csv files that users upload to an Amazon S3 bucket. The size of each S3 object is less than 100 MB. Which solution will meet these requirements MOST cost-effectively?
Community Votes
65% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The exam tests whether you size ETL compute to the data volume — the trap is assuming Spark/PySpark is always the 'proper' big-data answer when a fraction-of-a-DPU Python shell job is far cheaper for files under 100 MB.
Choosing the cheapest AWS service to run a daily ETL job over sub-100 MB CSV files uploaded to Amazon S3 comes down to matching compute overhead to data size. This page establishes that an AWS Glue Python shell job with pandas (D) beats EKS, EMR, and Glue PySpark on cost for this workload.
Picking option C (AWS Glue PySpark) because Apache Spark sounds more powerful, scalable, and 'production-grade' — but Spark's cluster provisioning and per-DPU-hour billing are unnecessary overhead when each input file is under 100 MB.
Community Discussion (14 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option D matches the AWS guidance that Glue Python shell jobs are intended for small-to-medium ETL workloads and can be launched on 1/16 DPU, so the effective per-run cost is a fraction of a full Spark job. At under 100 MB per file, pandas can load, clean, and rewrite each CSV in memory without distributed processing, so no Spark cluster is needed. Because the job runs once daily on modest data, minimizing DPU footprint matters more than horizontal scalability. As atu1789 put it, Python shell jobs are "a good fit for smaller-scale ETL tasks" exactly like these daily CSV uploads. Leo87656789 adds the decisive billing detail: "you can select the option \"1/16 DPU\" in the Job details for a Python Shell Job", which is cheaper per run than any Spark configuration.Why the Other Options Are Wrong
Option A (custom Python on Amazon EKS) forces you to build, patch, and pay for container infrastructure and worker nodes that sit idle between daily runs — the highest operational and monetary cost of all four. Option B (PySpark on Amazon EMR) means provisioning a cluster (or paying for transient cluster spin-up) for sub-100 MB files, again far more capacity than the job needs. Option C (AWS Glue PySpark) is serverless and legitimate for large datasets, but as khchan123 noted, "C may be overkill for processing small.csv files (less than 100 MB each)" because Spark's per-DPU-hour cost and startup overhead are not justified at this scale. The pricing quotes from halogi and lucas_rfsb compare raw per-DPU rates without factoring in that the shell job consumes only 1/16 of a DPU per run, so the actual invoice is lower for D.Community Comment Notes
The vote split (D 65 vs C 35) tracks a pricing misunderstanding: halogi and lucas_rfsb cite the official Glue pricing page showing shell jobs at $0.44 per DPU-hour versus PySpark at $0.29 per DPU-hour on flexible execution, concluding Spark is cheaper. That comparison ignores the DPU count actually consumed — Leo87656789's point about the 1/16 DPU setting resolves it in favor of D. cloudata reinforces the fit argument by linking the AWS Glue best-practices whitepaper's note that Python shell jobs suit smaller datasets, and khchan123 echoes the overkill concern for C. pypelyncar fairly calls it "50/50" on approach but still leans pandas, while not one commenter defends the EKS or EMR options.Official Reference
Exam Strategy
When a DEA-C01 question stresses 'MOST cost-effectively' and gives a small data size (under 100 MB, low file counts, daily batches), look for the smallest managed compute that fits — here the Glue Python shell job — and reject answers that require clusters (EKS, EMR) or distributed Spark. Always check whether a service's minimum billing unit (1/16 DPU vs 1 DPU) changes the real cost, not just the headline per-DPU rate.
Frequently Asked Questions
Why is an AWS Glue Python shell job cheaper than a Glue PySpark job here?
The shell job can run on 1/16 DPU and pandas handles sub-100 MB CSVs in memory, while a PySpark job bills full DPUs plus Spark startup overhead the small files never need.
When would option C, an AWS Glue PySpark job, actually be the better choice?
When files or volumes are large enough that data cannot fit in a single node's memory or the transform must be distributed across a Spark cluster; that scalability is wasted on under-100 MB daily CSVs.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →