How to Make Amazon Athena Queries on CSV Files Run Faster?
A data engineer needs Amazon Athena queries to finish faster. The data engineer notices that all the files the Athena queries use are currently stored in uncompressed .csv format. The data engineer also notices that users perform most queries by selecting a specific column. Which solution will MOST speed up the Athena query performance?
Community Votes
100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
It tests whether you know Athena charges and performs by data scanned, so a columnar format (Parquet) beats merely compressing the existing row-based CSV format — the trap is picking gzip or Snappy on .csv, which still reads every column.
This question asks which change will MOST speed up Amazon Athena queries whose source files are uncompressed .csv and whose users filter on a single column. The answer is to convert the data to Apache Parquet with Snappy compression, because columnar storage plus compression cuts the bytes Athena must scan.
Choosing gzip or Snappy compression on the .csv files (options B or D). Compression shrinks file size, but .csv stays row-oriented, so Athena still reads every column of every row; only a columnar format like Parquet enables column pruning and predicate pushdown.
Community Discussion (9 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Athena is a Presto/Trino-based engine that pays for and performs on the amount of data scanned, so the biggest win comes from reducing the columns and row groups touched per query. Apache Parquet is a columnar format, and the question states users "perform most queries by selecting a specific column," which is exactly the access pattern Parquet is optimized for through column pruning, predicate pushdown and per-column statistics. Layering Snappy on top of Parquet further shrinks the bytes read while remaining splittable and fast to decompress, which is why AWS's own tuning guidance lists converting to a columnar format with Snappy as the top Athena performance tip. Together, Parquet plus Snappy delivers the largest reduction in scanned data for selective, single-column queries.Why the Other Options Are Wrong
Option A moves.csv to JSON and adds Snappy, but JSON is a row-based text format that is typically larger than CSV, so it can make things worse rather than better. Option B compresses the existing.csv files with Snappy, which is a genuine improvement over plain CSV but leaves the row-oriented layout intact, so every column is still read for a one-column query. Option D uses gzip on.csv, which achieves similar row-format compression plus the downside that gzip files are not splittable in the same way, limiting parallelism. None of B, C and D alternatives preserve CSV while changing compression alone can match Parquet's columnar pruning gains, which is why C dominates.Community Comment Notes
Consensus in the comments is unanimous behind option C, with milofficial joking that "If the exam would only have these kinds of questions everyone would be blessed." TonyStark0122 spells out the mechanism, noting Parquet is a columnar format that enables column pruning and predicate pushdown — precisely the hint embedded in the question's single-column query pattern. wa212 and GabrielSGoncalves both point to the AWS Big Data blog on the top 10 performance tuning tips for Amazon Athena, and k350Secops and d8945a1 repeat that Parquet with Snappy gives the most significant gain because queries select a specific column.Official Reference
Exam Strategy
When an Athena question mentions slow queries over .csv, .json or other text files, look for the answer that changes the file format to a columnar one (Parquet or ORC) combined with Snappy compression — compression alone is never the 'MOST' improvement. Also treat phrases like 'queries select a specific column' as a direct hint toward columnar storage and predicate pushdown.
Frequently Asked Questions
Why isn't compressing the .csv files with Snappy or gzip enough for Athena?
Compression only shrinks bytes on disk; CSV is still row-based, so Athena reads every column for a single-column filter. Parquet adds column pruning plus predicate pushdown, cutting scanned data far more.
Why is JSON with Snappy worse than Parquet with Snappy here?
JSON is row-oriented text and usually larger than CSV, so it keeps the same full-row scan pattern. Only the columnar layout of Parquet lets Athena read just the selected column.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →