Which SageMaker Inference Option Fits Large Archived Datasets?
A company is building an ML model to analyze archived data. The company must perform inference on large datasets that are multiple GBs in size. The company does not need to access the model predictions immediately. Which Amazon SageMaker inference option will meet these requirements?
Community Votes
78% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
This question tests the distinction between Batch Transform (offline, bulk, no endpoint required) and Asynchronous Inference (near-real-time with queues), a common trap for candidates who conflate 'not immediate' with 'asynchronous'.
For offline analysis of multi-GB archived data where predictions are not needed immediately, Amazon SageMaker Batch Transform is the most cost-effective and scalable inference option. Community consensus strongly favors Batch Transform over Asynchronous Inference for bulk, non-real-time workloads.
Many candidates choose Asynchronous Inference (D) because the prompt says predictions are not needed immediately, but Asynchronous Inference is designed for streaming or near-real-time workloads with payloads up to 1 GB, not multi-GB archived datasets processed in bulk.
Community Discussion (6 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why Batch Transform Is Correct
Amazon SageMaker Batch Transform is purpose-built for offline, bulk inference on large datasets stored in Amazon S3. Key characteristics that match the scenario:
- No persistent endpoint required – you launch a transform job, it spins up the model, processes the entire dataset, writes results back to S3, and tears down. This makes it extremely cost-effective for archived data.
- Handles multi-GB (even TB-scale) datasets by automatically splitting input data into manageable batches.
- No latency requirement – results are available once the job completes, which perfectly fits "does not need to access the model predictions immediately."
Why the Other Options Are Wrong
- B. Real-time inference – Provides a persistent endpoint for low-latency, on-demand predictions (milliseconds). It is overkill and far too expensive for bulk archived data.
- C. Serverless inference – Designed for sporadic, unpredictable traffic with payloads typically under 6 MB. It cannot handle multi-GB datasets.
- D. Asynchronous inference – While it also does not require immediate results, it is intended for near-real-time workloads with large payloads (up to 1 GB per invocation) that are queued and processed with SLAs in minutes. It still requires an endpoint and is not optimized for purely offline, bulk batch jobs on archived data. Comment [4] reflects the common confusion, but AWS documentation clearly positions Batch Transform as the correct choice for offline bulk processing.
Key Takeaway
When the prompt mentions archived data, multiple GBs, and no need for immediate predictions, always default to Batch Transform. Reserve Asynchronous Inference for workloads that still need an endpoint, have SLAs measured in minutes, and involve streaming or large individual payloads.
Official Reference
Exam Strategy
On the exam, immediately map keywords to SageMaker inference types: 'archived/bulk/offline + multi-GB + no immediate need' = Batch Transform; 'large payload + minutes SLA + endpoint needed' = Asynchronous; 'millisecond latency' = Real-time; 'sporadic traffic + small payload' = Serverless. Eliminating options by payload size and latency requirement will quickly isolate the correct answer.
Related Analysis
Practice All AIF-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full AIF-C01 Practice Test →