Most Cost-Effective AWS Glue ETL for Small Daily CSV Files?

Answer Correct answer: D — Use an AWS Glue Python shell job with pandas, because sub-100 MB daily CSV files are transformed most cheaply on a 1/16 DPU shell job.

A data engineer needs to build an extract, transform, and load (ETL) job. The ETL job will process daily incoming .csv files that users upload to an Amazon S3 bucket. The size of each S3 object is less than 100 MB. Which solution will meet these requirements MOST cost-effectively?

  1. Write a custom Python application. Host the application on an Amazon Elastic Kubernetes Service (Amazon EKS) cluster.
  2. Write a PySpark ETL script. Host the script on an Amazon EMR cluster.
  3. Write an AWS Glue PySpark job. Use Apache Spark to transform the data.
  4. Write an AWS Glue Python shell job. Use pandas to transform the data. Correct Answer

Community Votes

D
65%
C
35%

65% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The exam tests whether you size ETL compute to the data volume — the trap is assuming Spark/PySpark is always the 'proper' big-data answer when a fraction-of-a-DPU Python shell job is far cheaper for files under 100 MB.

Choosing the cheapest AWS service to run a daily ETL job over sub-100 MB CSV files uploaded to Amazon S3 comes down to matching compute overhead to data size. This page establishes that an AWS Glue Python shell job with pandas (D) beats EKS, EMR, and Glue PySpark on cost for this workload.

Picking option C (AWS Glue PySpark) because Apache Spark sounds more powerful, scalable, and 'production-grade' — but Spark's cluster provisioning and per-DPU-hour billing are unnecessary overhead when each input file is under 100 MB.

Community Discussion (14 comments)

halogi 👍 10 Selected: C
AWS Glue Python Shell Job is billed $0.44 per DPU-Hour for each job AWS Glue PySpark is billed $0.29 per DPU-Hour for each job with flexible execution and $0.44 per DPU-Hour for each job with standard execution Source: https://aws.amazon.com/glue/pricing/
atu1789 👍 9 Selected: D
Option D: Write an AWS Glue Python shell job and use pandas to transform the data, is the most cost-effective solution for the described scenario. AWS Glue’s Python shell jobs are a good fit for smaller-scale ETL tasks, especially when dealing with .csv files that are less than 100 MB each. The use of pandas, a powerful and efficient data manipulation library in Python, makes it an ideal tool for processing and transforming these types of files. This approach avoids the overhead and additional costs associated with more complex solutions like Amazon EKS or EMR, which are generally more suited for larger-scale, more complex data processing tasks. Given the requirements – processing daily incoming small-sized .csv files – this solution provides the necessary functionality with minimal resources, aligning well with the goal of cost-effectiveness.
YUICH 👍 2 Selected: D
It is important not to compare just the “price per DPU hour,” but to consider the total cost by factoring in overhead for job startup, minimum DPU count, execution time, and data volume. For a relatively lightweight workload—such as processing approximately 100 MB of CSV files on a daily basis—option (D), using an AWS Glue Python shell job, is the most cost-effective choice.
LR2023 👍 3 Selected: D
going with D https://docs.aws.amazon.com/whitepapers/latest/aws-glue-best-practices-build-performant-data-pipeline/additional-considerations.html
pypelyncar 👍 7 Selected: D
good candidate to be (2 options) for real, either spark and py have similar approaches. I would go with Pandas, although... 50/50.. it could be Spark. I hope not to find this question in the exam
VerRi 👍 3 Selected: C
PySpark with Spark(Flexible Execution): $0.29/hr for 1 DPU PySpark with Spark(Standard Execution): $0.44/hr for 1 DPU Python Shell with Pandas: $0.44/hr for 1 DPU
cloudata 👍 6 Selected: D
Python Shell is cheaper and can handle small to medium tasks. https://docs.aws.amazon.com/whitepapers/latest/aws-glue-best-practices-build-performant-data-pipeline/additional-considerations.html
chakka90 👍 3
D. Because the pyspark is still being the cheap you have to use minimum of 2 DPU. Which would increase the cost anyway so, i feel that d should be correct
khchan123 👍 4 Selected: D
D. While AWS Glue PySpark jobs are scalable and suitable for large workloads, C may be overkill for processing small .csv files (less than 100 MB each). The overhead of using Apache Spark may not be cost-effective for this specific use case.
Leo87656789 👍 4 Selected: D
Option D: Even though the Python Shell Job is more expensive on a DPU-Hour basis, you can select the option "1/16 DPU" in the Job details for a Python Shell Job, which is definetly cheaper than a Pyspark job.
lucas_rfsb 👍 6 Selected: C
AWS Glue Python Shell Job is billed $0.44 per DPU-Hour for each job AWS Glue PySpark is billed $0.29 per DPU-Hour for each job with flexible execution and $0.44 per DPU-Hour for each job with standard execution Source: https://aws.amazon.com/glue/pricing/
[Removed] 👍 5 Selected: D
https://medium.com/@navneetsamarth/reduce-aws-cost-using-glue-python-shell-jobs-70a955d4359f#:~:text=The%20cheapest%20Glue%20Spark%20ETL,1%2F16th%20of%20a%20DPU.&text=This%20can%20result%20in%20massive,just%20a%20better%20design%20overall!
GiorgioGss 👍 4 Selected: D
D is more cheaper than C. Not so scalable but is cheaper...
rralucard_ 👍 5 Selected: C
AWS Glue is a fully managed ETL service, which means you don't need to manage infrastructure, and it automatically scales to handle your data processing needs. This reduces operational overhead and cost. PySpark, as a part of AWS Glue, is a powerful and widely-used framework for distributed data processing, and it's well-suited for handling data transformations on a large scale.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option D matches the AWS guidance that Glue Python shell jobs are intended for small-to-medium ETL workloads and can be launched on 1/16 DPU, so the effective per-run cost is a fraction of a full Spark job. At under 100 MB per file, pandas can load, clean, and rewrite each CSV in memory without distributed processing, so no Spark cluster is needed. Because the job runs once daily on modest data, minimizing DPU footprint matters more than horizontal scalability. As atu1789 put it, Python shell jobs are "a good fit for smaller-scale ETL tasks" exactly like these daily CSV uploads. Leo87656789 adds the decisive billing detail: "you can select the option \"1/16 DPU\" in the Job details for a Python Shell Job", which is cheaper per run than any Spark configuration.

Why the Other Options Are Wrong

Option A (custom Python on Amazon EKS) forces you to build, patch, and pay for container infrastructure and worker nodes that sit idle between daily runs — the highest operational and monetary cost of all four. Option B (PySpark on Amazon EMR) means provisioning a cluster (or paying for transient cluster spin-up) for sub-100 MB files, again far more capacity than the job needs. Option C (AWS Glue PySpark) is serverless and legitimate for large datasets, but as khchan123 noted, "C may be overkill for processing small.csv files (less than 100 MB each)" because Spark's per-DPU-hour cost and startup overhead are not justified at this scale. The pricing quotes from halogi and lucas_rfsb compare raw per-DPU rates without factoring in that the shell job consumes only 1/16 of a DPU per run, so the actual invoice is lower for D.

Community Comment Notes

The vote split (D 65 vs C 35) tracks a pricing misunderstanding: halogi and lucas_rfsb cite the official Glue pricing page showing shell jobs at $0.44 per DPU-hour versus PySpark at $0.29 per DPU-hour on flexible execution, concluding Spark is cheaper. That comparison ignores the DPU count actually consumed — Leo87656789's point about the 1/16 DPU setting resolves it in favor of D. cloudata reinforces the fit argument by linking the AWS Glue best-practices whitepaper's note that Python shell jobs suit smaller datasets, and khchan123 echoes the overkill concern for C. pypelyncar fairly calls it "50/50" on approach but still leans pandas, while not one commenter defends the EKS or EMR options.

Official Reference

Exam Strategy

When a DEA-C01 question stresses 'MOST cost-effectively' and gives a small data size (under 100 MB, low file counts, daily batches), look for the smallest managed compute that fits — here the Glue Python shell job — and reject answers that require clusters (EKS, EMR) or distributed Spark. Always check whether a service's minimum billing unit (1/16 DPU vs 1 DPU) changes the real cost, not just the headline per-DPU rate.

Frequently Asked Questions

Why is an AWS Glue Python shell job cheaper than a Glue PySpark job here?

The shell job can run on 1/16 DPU and pandas handles sub-100 MB CSVs in memory, while a PySpark job bills full DPUs plus Spark startup overhead the small files never need.

When would option C, an AWS Glue PySpark job, actually be the better choice?

When files or volumes are large enough that data cannot fit in a single node's memory or the transform must be distributed across a Spark cluster; that scalability is wasted on under-100 MB daily CSVs.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide