Cost-Effective Querying of Compressed Data for Audits

Answer Correct answer: A — Store the data in Amazon Glacier Flexible Retrieval and use Amazon S3 Glacier Select to query the data efficiently.

An insurance company stores transaction data that the company compressed with gzip. The company needs to query the transaction data for occasional audits. Which solution will meet this requirement in the MOST cost-effective way?

  1. Store the data in Amazon Glacier Flexible Retrieval. Use Amazon S3 Glacier Select to query the data. Correct Answer
  2. Store the data in Amazon S3. Use Amazon S3 Select to query the data.
  3. Store the data in Amazon S3. Use Amazon Athena to query the data.
  4. Store the data in Amazon Glacier Instant Retrieval. Use Amazon Athena to query the data.

Community Votes

B
56%
A
44%

56% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The core concept tested is optimizing for 'occasional' access patterns by selecting a low-cost storage class (Glacier) that still supports direct querying capabilities, avoiding the higher per-query costs or egress fees associated with standard S3 services like Athena or S3 Select.

This question evaluates the most cost-effective solution for querying gzip-compressed transaction data stored in AWS, focusing on balancing storage costs with query performance and pricing. The correct approach leverages Amazon S3 Glacier Flexible Retrieval combined with Amazon S3 Glacier Select to minimize expenses for occasional audits.

Many learners choose Option B (S3 Standard with S3 Select) because they prioritize fast retrieval speeds over long-term storage costs, failing to recognize that 'occasional audits' justify the slower retrieval times of Glacier in exchange for significant savings.

Community Discussion (23 comments)

tgv 👍 9 Selected: A
Actually, I think A makes more sense.
AM027 👍 1 Selected: A
cold data stored in Glacier can be easily queried within minutes.
YUICH 👍 1 Selected: B
For workloads with low access frequency where you only need to query data occasionally (for example, during audits), option (A)—S3 Glacier Flexible Retrieval combined with S3 Glacier Select—provides the most cost-effective solution.
div_div 👍 1 Selected: C
Transaction Data Refers To Data Which Are Updating Frequently and To Query That Data occasionally Means It Can Be Query At Any Time (In Question Time Is Not Define). So We Can't Take Risk For Customer To Wait For Hours To Get The Result And The Best Way To Query The Data On Top of The S3 Bucket We Can Use Athena.
BigMrT 👍 1 Selected: B
Glacier Select incurs higher costs compared to S3 Select.
ctndba 👍 1
You cannot use S3 Select on S3 Glacier Flexible Retrieval storage class. So answer is B, based on given options.
mohamedTR 👍 2 Selected: B
B is the more cost-effective solution for occasional audits. It allows for easier access to the data without incurring high retrieval costs
manig 👍 1
gzip compressed data querying -> s3 select -Answer B
LR2023 👍 2 Selected: A
https://aws.amazon.com/blogs/aws/s3-glacier-select/ option B is not cost effective as it is stored in standard S3
PashoQ 👍 2 Selected: A
Occasional audits, so go for S3 glacier select
cas_tori 👍 4 Selected: B
this is B
IanJang 👍 1
IT is A
mns0173 👍 2
Glacier is an expensive option in cases when you need to access data occasionally
lenneth39 👍 1 Selected: C
I am not sure whether to go for B or C. Can anyone comment on this? B: No problem, but not available if Parquet is Gzip compressed. But the problem statement doesn't say Parquet is Gzip compressed. C: Correct if Parquet is Gzip compressed, but B is more cost-effective if csv or json is Gzip compressed
andrologin 👍 3 Selected: B
I think the solution is either B or D but I would go with B because they mentioned storing the data in gzip and not parquet which is optimised for Athena queries
4bc91ae 👍 2
there is no such thing as Glacier Flexible Retrieval, so its no A . its either B or D and most likely its D for the cost
bakarys 👍 2 Selected: B
B. Store the data in Amazon S3. Use Amazon S3 Select to query the data. Amazon S3 is a cost-effective object storage service, and S3 Select allows you to retrieve only a subset of data from an object by using simple SQL expressions. S3 Select works on objects stored in CSV, JSON, or Apache Parquet format. It also supports GZIP and BZIP2 compression formats, which makes it suitable for the given scenario where the data is compressed with gzip. While Amazon Athena is a powerful query service, it can be more expensive than S3 Select for occasional queries. Amazon Glacier and Glacier Select are designed for long-term archival storage and not for frequent access or queries, which might not be suitable for occasional audits. Therefore, option B is the most cost-effective choice for this scenario.
FunkyFresco 👍 2 Selected: B
ill go with B, because. of cost to query
bakarys 👍 2 Selected: B
B. Store the data in Amazon S3. Use Amazon S3 Select to query the data. Amazon S3 is a cost-effective storage service, and S3 Select allows you to retrieve only a subset of data from an object by using simple SQL expressions. S3 Select works on objects stored in CSV, JSON, or Apache Parquet format. It also supports GZIP compression, which is the format used by the company. This makes it a cost-effective solution for occasional queries needed for audits.
Alagong 👍 3 Selected: B
Option B (Amazon S3 with S3 Select) is generally more cost-effective and operationally efficient for occasional audits of gzip-compressed data. It provides faster access to data and lower querying costs, which are critical factors for ad-hoc and timely data retrievals. While Option A (Amazon Glacier Flexible Retrieval with S3 Glacier Select) offers cheaper storage, its longer retrieval times and potential higher querying costs make it less suitable for use cases requiring timely access to data.
HunkyBunky 👍 3
Looks like that A - fit better in question requirements
GHill1982 👍 3 Selected: B
On the assumptions that querying the audit data is time sensitive and the transaction data is compressed into a single object I would go with using S3 and S3 select to query the data.
artworkad 👍 4 Selected: A
B and C are not cost effective. A is more cost effective than D. I go with A.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option A is the correct answer because it addresses both the storage format (gzip) and the access pattern (occasional audits). Amazon S3 Glacier Flexible Retrieval provides the lowest storage cost among options that support querying. Amazon S3 Glacier Select allows you to filter data using SQL queries directly within the compressed files without retrieving the entire object first, which avoids high data retrieval and processing costs associated with services like Athena or full restores.

Why the Other Options Are Wrong

Option B (S3 Standard) incurs significantly higher monthly storage costs compared to Glacier, making it less cost-effective for data that is rarely accessed. Option C (Athena on S3 Standard) adds query processing costs on top of standard storage fees, which is inefficient for occasional use cases. Option D is incorrect because Amazon Athena cannot natively query data stored in Glacier; it requires data to be in S3 Standard or compatible classes, and attempting to do so would incur massive retrieval fees or fail entirely.

Community Comment Notes

Several community members, such as tgv and artworkad, correctly identified that Option A is more suitable than B or C due to cost constraints for infrequent access. User LR2023 highlighted that storing data in S3 Standard (Option B) is not cost-effective for this scenario, reinforcing the need for Glacier. While some users argued for speed (Option B), the question explicitly prioritizes 'MOST cost-effective', validating the choice of Glacier for archival-like audit data.

Exam Strategy

When questions emphasize 'cost-effective' for data that is 'infrequently' or 'occasionally' accessed, always consider Amazon S3 Glacier Flexible Retrieval or Deep Archive. Ensure the service selected supports querying capabilities if the requirement involves analyzing the data without full restoration.

Frequently Asked Questions

Why is S3 Select not enough for cost-effectiveness?

S3 Select reduces data scanned during queries, but storing large volumes of historical audit data in S3 Standard incurs high monthly storage fees compared to Glacier.

Can Athena query gzip files in Glacier?

No, Athena operates on data in S3 Standard or similar tiers. It cannot directly query objects in Glacier without first restoring them, which defeats the purpose of cost optimization.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide