Store VPC flow logs in Apache Parquet with hourly partitions for Athena
A company deploys workloads in multiple AWS accounts. Each account has a VPC with VPC flow logs published in text log format to a centralized Amazon S3 bucket. Each log file is compressed with gzip compression. The company must retain the log files indefinitely. A security engineer occasionally analyzes the logs by using Amazon Athena to query the VPC flow logs. The query performance is degrading over time as the number of ingested logs is growing. A solutions architect must improve the performance of the log analysis and reduce the storage space that the VPC flow logs use. Which solution will meet these requirements with the LARGEST performance improvement?
Community Votes
100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
Athena reads only the columns referenced by a query, and Parquet is a columnar format with per-column compression, so converting the flow logs to Parquet reduces both the bytes scanned and the storage used, and hourly partitioning lets Athena prune entire prefixes from the query's scope.
VPC flow logs from many accounts are published as gzip-compressed text files to a centralized S3 bucket, retained indefinitely, and occasionally queried with Amazon Athena. Query performance degrades as the number of ingested logs grows, and both query speed and storage footprint need improvement.
Re-compressing the files with bzip2. Query performance in Athena is dominated by how much data it must read and in what layout, not by the compression codec alone, so swapping gzip for bzip2 on the same text layout yields a marginal storage gain and requires a Lambda function on every object creation. Recompressing to Intelligent-Tiering likewise does not change the queryable layout.
Community Discussion (6 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Athena performance is governed by the amount of data it must scan and the columnar layout of that data. Changing the VPC flow log configuration to publish in Apache Parquet format means Athena reads only the columns a query actually references instead of parsing whole text lines, which is the single largest improvement available. Parquet is also compressed per column, so the same data occupies less S3 storage, which addresses the storage reduction requirement at the same time. Specifying hourly partitions adds a prefix structure that lets Athena prune whole partitions from the query, so a query over a recent window does not have to enumerate objects across the entire retained history, which is what makes performance degrade as ingestion grows.Why the Other Options Are Wrong
A: A Lambda function that decompresses gzip and recompresses with bzip2 changes the codec but keeps the same text row layout, so Athena still parses the same amount of data with the same column handling, and it adds a Lambda function, permissions, and a per-object invocation cost on every upload. B: S3 Transfer Acceleration affects upload and download speed, not Athena query performance, and moving files to Intelligent-Tiering changes the storage class rather than the file format, so queries still scan the same text layout. D: A new workgroup without data usage limits and Athena engine version 2 can change cost controls and some query features, but neither changes the underlying file format or partition layout, so the scan cost and degradation over time are unchanged.Community Comment Notes
The community voted 100 to 0 for C, and the reasoning was consistent across the top comments: Parquet is designed for efficient querying because Athena can read only the columns it needs, and hourly partitions let Athena skip irrelevant data. Commenters linked the AWS announcement about VPC flow logs in Parquet with hive-prefixed partitioned files, which is the feature that delivers exactly this combination.Official Reference
Related Analysis
Practice All SAP-C02 Questions
Access 85 questions with complete answers and detailed explanations.
View Full SAP-C02 Practice Test →