How to Make Amazon Athena Queries on CSV Files Run Faster?

Answer Correct answer: C — Convert the uncompressed .csv data to Apache Parquet with Snappy compression so Athena can prune columns and scan far fewer bytes per query.

A data engineer needs Amazon Athena queries to finish faster. The data engineer notices that all the files the Athena queries use are currently stored in uncompressed .csv format. The data engineer also notices that users perform most queries by selecting a specific column. Which solution will MOST speed up the Athena query performance?

  1. Change the data format from .csv to JSON format. Apply Snappy compression.
  2. Compress the .csv files by using Snappy compression.
  3. Change the data format from .csv to Apache Parquet. Apply Snappy compression. Correct Answer
  4. Compress the .csv files by using gzip compression.

Community Votes

C
100%

100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

It tests whether you know Athena charges and performs by data scanned, so a columnar format (Parquet) beats merely compressing the existing row-based CSV format — the trap is picking gzip or Snappy on .csv, which still reads every column.

This question asks which change will MOST speed up Amazon Athena queries whose source files are uncompressed .csv and whose users filter on a single column. The answer is to convert the data to Apache Parquet with Snappy compression, because columnar storage plus compression cuts the bytes Athena must scan.

Choosing gzip or Snappy compression on the .csv files (options B or D). Compression shrinks file size, but .csv stays row-oriented, so Athena still reads every column of every row; only a columnar format like Parquet enables column pruning and predicate pushdown.

Community Discussion (9 comments)

milofficial 👍 11 Selected: C
If the exam would only have these kinds of questions everyone would be blessed
TonyStark0122 👍 6
C. Change the data format from .csv to Apache Parquet. Apply Snappy compression. Explanation: Apache Parquet is a columnar storage format optimized for analytical queries. It is highly efficient for query performance, especially when queries involve selecting specific columns, as it allows for column pruning and predicate pushdown optimizations.
Scotty_Nguyen 👍 1 Selected: C
C is correct
GabrielSGoncalves 👍 1 Selected: C
C is the way to do It based on best practices recommended by AWS (https://aws.amazon.com/pt/blogs/big-data/top-10-performance-tuning-tips-for-amazon-athena/)
hnk 👍 1 Selected: C
C is correct
k350Secops 👍 1 Selected: C
switching to Apache Parquet format with Snappy compression offers the most significant improvement in Athena query performance, especially for queries that select specific columns
d8945a1 👍 1 Selected: C
Parquet is columnar storage and the question specifies that users performs most queries by selecting a specific column.
wa212 👍 2 Selected: C
https://aws.amazon.com/jp/blogs/news/top-10-performance-tuning-tips-for-amazon-athena/
Alcee 👍 1
C easy

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Athena is a Presto/Trino-based engine that pays for and performs on the amount of data scanned, so the biggest win comes from reducing the columns and row groups touched per query. Apache Parquet is a columnar format, and the question states users "perform most queries by selecting a specific column," which is exactly the access pattern Parquet is optimized for through column pruning, predicate pushdown and per-column statistics. Layering Snappy on top of Parquet further shrinks the bytes read while remaining splittable and fast to decompress, which is why AWS's own tuning guidance lists converting to a columnar format with Snappy as the top Athena performance tip. Together, Parquet plus Snappy delivers the largest reduction in scanned data for selective, single-column queries.

Why the Other Options Are Wrong

Option A moves.csv to JSON and adds Snappy, but JSON is a row-based text format that is typically larger than CSV, so it can make things worse rather than better. Option B compresses the existing.csv files with Snappy, which is a genuine improvement over plain CSV but leaves the row-oriented layout intact, so every column is still read for a one-column query. Option D uses gzip on.csv, which achieves similar row-format compression plus the downside that gzip files are not splittable in the same way, limiting parallelism. None of B, C and D alternatives preserve CSV while changing compression alone can match Parquet's columnar pruning gains, which is why C dominates.

Community Comment Notes

Consensus in the comments is unanimous behind option C, with milofficial joking that "If the exam would only have these kinds of questions everyone would be blessed." TonyStark0122 spells out the mechanism, noting Parquet is a columnar format that enables column pruning and predicate pushdown — precisely the hint embedded in the question's single-column query pattern. wa212 and GabrielSGoncalves both point to the AWS Big Data blog on the top 10 performance tuning tips for Amazon Athena, and k350Secops and d8945a1 repeat that Parquet with Snappy gives the most significant gain because queries select a specific column.

Official Reference

Exam Strategy

When an Athena question mentions slow queries over .csv, .json or other text files, look for the answer that changes the file format to a columnar one (Parquet or ORC) combined with Snappy compression — compression alone is never the 'MOST' improvement. Also treat phrases like 'queries select a specific column' as a direct hint toward columnar storage and predicate pushdown.

Frequently Asked Questions

Why isn't compressing the .csv files with Snappy or gzip enough for Athena?

Compression only shrinks bytes on disk; CSV is still row-based, so Athena reads every column for a single-column filter. Parquet adds column pruning plus predicate pushdown, cutting scanned data far more.

Why is JSON with Snappy worse than Parquet with Snappy here?

JSON is row-oriented text and usually larger than CSV, so it keeps the same full-row scan pattern. Only the columnar layout of Parquet lets Athena read just the selected column.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide