Archive zipped image batches to S3 Glacier Deep Archive with a DataSync Hyper-V agent
A company has an application that analyzes and stores image data on premises. The application receives millions of new image files every day. Files are an average of 1 MB in size. The files are analyzed in batches of 1 GB. When the application analyzes a batch, the application zips the images together. The application then archives the images as a single file in an on-premises NFS server for long-term storage. The company has a Microsoft Hyper-V environment on premises and has compute capacity available. The company does not have storage capacity and wants to archive the images on AWS. The company needs the ability to retrieve archived data within 1 week of a request. The company has a 10 Gbps AWS Direct Connect connection between its on-premises data center and AWS. The company needs to set bandwidth limits and schedule archived images to be copied to AWS during non-business hours. Which solution will meet these requirements MOST cost-effectively?
Community Votes
100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The requirement is to retrieve within one week, which rules out the standard Glacier retrieval tiers, and S3 Glacier Deep Archive is the lowest-cost class that still meets it, while a DataSync agent on an existing Hyper-V host avoids buying any new compute.
An on-premises application zips images into single files and archives them to an NFS server. The company has a Hyper-V environment, spare compute capacity, no storage capacity, a 10 Gbps Direct Connect link, and needs archived data retrievable within one week of a request, with bandwidth limits and off-hours scheduling.
Using S3 Glacier Instant Retrieval because the data must be quickly accessible. Instant Retrieval is the fastest and most expensive Glacier class, and paying that price for data that is requested at most once a week defeats the most cost-effective requirement. Storing in S3 Standard first and transitioning a day later also leaves a day of Standard-tier storage for no benefit.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The data must be retrievable within one week, which is the constraint that determines the storage class. S3 Glacier Deep Archive is the lowest-cost S3 storage class, and its standard retrieval completes within twelve hours, comfortably inside the one-week requirement, so it is the cheapest tier that qualifies. DataSync is a managed transfer service that schedules the copy, applies bandwidth limits, and runs on a schedule, which satisfies the non-business-hours and bandwidth requirements without custom scripting. The agent is deployed as a Hyper-V virtual machine on the existing on-premises host, so the company uses compute it already has and the agent can read the NFS source directly, and the 10 Gbps Direct Connect link carries the traffic once the schedule window opens.Why the Other Options Are Wrong
A: Deploying a DataSync agent on a new GPU-based EC2 instance adds GPU instance cost for a data transfer task that needs no GPU, and S3 Glacier Instant Retrieval is the most expensive of the Glacier classes, which is the wrong direction for a most cost-effective requirement that only needs retrieval within a week. C: Storing to S3 Standard and then transitioning to Deep Archive after one day pays a full day of Standard storage for data that is immediately archive-bound, adding cost and delay for no benefit. D: A tape gateway with automatic tape creation is designed to write to physical tape cartridges for edge and backup scenarios, and it introduces a tape handling and rotation process, which is significantly more operational overhead than a managed DataSync copy that the company can schedule and forget.Community Comment Notes
The community voted 100 to 0 for B, and the reasoning was tightly focused on the retrieval window: a commenter stated that option A is out because Glacier Instant Retrieval is for millisecond access, B is correct because it goes directly to Deep Archive, and C needlessly keeps the data in S3 Standard for a day. Another commenter noted the D option is an awkward use case, and two linked the AWS blog on protecting archives with DataSync and S3 Glacier.Official Reference
Related Analysis
Practice All SAP-C02 Questions
Access 85 questions with complete answers and detailed explanations.
View Full SAP-C02 Practice Test →