Store VPC flow logs in Apache Parquet with hourly partitions for Athena

Answer Correct answer: C — Change the VPC flow log configuration to Apache Parquet format with hourly partitions.

A company deploys workloads in multiple AWS accounts. Each account has a VPC with VPC flow logs published in text log format to a centralized Amazon S3 bucket. Each log file is compressed with gzip compression. The company must retain the log files indefinitely. A security engineer occasionally analyzes the logs by using Amazon Athena to query the VPC flow logs. The query performance is degrading over time as the number of ingested logs is growing. A solutions architect must improve the performance of the log analysis and reduce the storage space that the VPC flow logs use. Which solution will meet these requirements with the LARGEST performance improvement?

  1. Create an AWS Lambda function to decompress the gzip files and to compress the files with bzip2 compression. Subscribe the Lambda function to an s3:ObjectCreated:Put S3 event notification for the S3 bucket.
  2. Enable S3 Transfer Acceleration for the S3 bucket. Create an S3 Lifecycle configuration to move files to the S3 Intelligent-Tiering storage class as soon as the files are uploaded.
  3. Update the VPC flow log configuration to store the files in Apache Parquet format. Specify hourly partitions for the log files. Correct Answer
  4. Create a new Athena workgroup without data usage control limits. Use Athena engine version 2.

Community Votes

C
100%

100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Athena reads only the columns referenced by a query, and Parquet is a columnar format with per-column compression, so converting the flow logs to Parquet reduces both the bytes scanned and the storage used, and hourly partitioning lets Athena prune entire prefixes from the query's scope.

VPC flow logs from many accounts are published as gzip-compressed text files to a centralized S3 bucket, retained indefinitely, and occasionally queried with Amazon Athena. Query performance degrades as the number of ingested logs grows, and both query speed and storage footprint need improvement.

Re-compressing the files with bzip2. Query performance in Athena is dominated by how much data it must read and in what layout, not by the compression codec alone, so swapping gzip for bzip2 on the same text layout yields a marginal storage gain and requires a Lambda function on every object creation. Recompressing to Intelligent-Tiering likewise does not change the queryable layout.

Community Discussion (6 comments)

AzureDP900 👍 2
C is correct -- The Parquet format is designed for efficient querying, and hourly partitions allow Athena to quickly scan through the logs. This combination will result in the largest performance improvement for log analysis compared to other options.
TonytheTiger 👍 4 Selected: C
Option C : https://aws.amazon.com/about-aws/whats-new/2021/10/amazon-vpc-flow-logs-parquet-hive-prefixes-partitioned-files/ New features to make it faster, easier and more cost efficient to store and run analytics on your Amazon VPC Flow Logs
pangchn 👍 4 Selected: C
C Using AWS Athena with parquet files is faster and cheaper than using other formats like CSV and JSON based file structures, according to AWS Athena pricing "compressing your data allows Athena to scan less data, and converting your data to columnar formats allows Athena to selectively read only required columns to process the data, which leads to cost savings and improved performance https://www.linkedin.com/pulse/aws-athena-parquet-vs-csv-ahmed-fayed/
lasithasilva709 👍 2 Selected: C
Apache Parquet is compressed, efficient columnar data representation https://parquet.apache.org/docs/overview/motivation/
Dgix 👍 1 Selected: C
C is correct.
CMMC 👍 1 Selected: C
Apache Parquet format to enable a highly optimized columnar storage format and partitioning by hour for improving the Athena query performance

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Athena performance is governed by the amount of data it must scan and the columnar layout of that data. Changing the VPC flow log configuration to publish in Apache Parquet format means Athena reads only the columns a query actually references instead of parsing whole text lines, which is the single largest improvement available. Parquet is also compressed per column, so the same data occupies less S3 storage, which addresses the storage reduction requirement at the same time. Specifying hourly partitions adds a prefix structure that lets Athena prune whole partitions from the query, so a query over a recent window does not have to enumerate objects across the entire retained history, which is what makes performance degrade as ingestion grows.

Why the Other Options Are Wrong

A: A Lambda function that decompresses gzip and recompresses with bzip2 changes the codec but keeps the same text row layout, so Athena still parses the same amount of data with the same column handling, and it adds a Lambda function, permissions, and a per-object invocation cost on every upload. B: S3 Transfer Acceleration affects upload and download speed, not Athena query performance, and moving files to Intelligent-Tiering changes the storage class rather than the file format, so queries still scan the same text layout. D: A new workgroup without data usage limits and Athena engine version 2 can change cost controls and some query features, but neither changes the underlying file format or partition layout, so the scan cost and degradation over time are unchanged.

Community Comment Notes

The community voted 100 to 0 for C, and the reasoning was consistent across the top comments: Parquet is designed for efficient querying because Athena can read only the columns it needs, and hourly partitions let Athena skip irrelevant data. Commenters linked the AWS announcement about VPC flow logs in Parquet with hive-prefixed partitioned files, which is the feature that delivers exactly this combination.

Official Reference

Related Analysis

Practice All SAP-C02 Questions

Access 85 questions with complete answers and detailed explanations.

View Full SAP-C02 Practice Test →

← Back to SAP-C02 Study Guide