Which SageMaker Inference Option Fits 1GB Payloads and 1-Hour Processing with Near Real-Time Latency?
A company uses Amazon SageMaker for its ML pipeline in a production environment. The company has large input data sizes up to 1 GB and processing times up to 1 hour. The company needs near real-time latency. Which SageMaker inference option meets these requirements?
Community Votes
79% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
This question tests your ability to match SageMaker inference options to specific payload size, processing time, and latency requirements, with the key trap being the phrase 'near real-time latency' which points directly to Asynchronous Inference rather than Real-time Inference.
Amazon SageMaker Asynchronous Inference is designed for large payloads (up to 1GB) and long processing times (up to 1 hour) while still offering near real-time latency. Community consensus strongly favors Asynchronous Inference (C) over Real-time Inference (A) for this scenario.
Many candidates incorrectly choose Real-time Inference (A) because they focus on the 'near real-time latency' requirement without considering that Real-time Inference cannot handle 1GB payloads or 1-hour processing times effectively.
Community Discussion (15 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Understanding Amazon SageMaker Inference Options
Amazon SageMaker offers four distinct inference options, each designed for specific use cases:
Real-time Inference (Option A) is designed for low-latency, high-throughput applications where predictions need to be made immediately. It's ideal for interactive applications requiring millisecond to second-level response times. However, Real-time Inference has practical limitations with very large payloads (typically limited to around 6MB for direct invocation) and cannot handle processing times approaching 1 hour.
Asynchronous Inference (Option C) is specifically designed for scenarios involving large payloads (up to 1GB) and long processing times (up to 1 hour) while still providing near real-time latency. As multiple community members correctly noted, this option queues incoming requests and processes them asynchronously, decoupling request submission from processing. This makes it perfect for the scenario described in the question.
Serverless Inference (Option B) is optimized for intermittent or unpredictable traffic patterns with low-latency requirements, but it's not designed for the large payload sizes and long processing times specified in this scenario.
Batch Transform (Option D) is meant for offline, large-scale processing without real-time needs. It's not suitable when near real-time latency is required.
Why Asynchronous Inference is Correct
The question presents three critical requirements: 1. Large input data sizes up to 1GB 2. Processing times up to 1 hour 3. Near real-time latency
Only Asynchronous Inference satisfies all three requirements simultaneously. As community member [14] correctly quoted from AWS documentation: "Amazon SageMaker Asynchronous Inference is a capability in SageMaker that queues incoming requests and processes them asynchronously. This option is ideal for requests with large payload sizes (up to 1GB), long processing times (up to one hour), and near real-time latency requirements."
Common Misconception
The trap in this question is the phrase "near real-time latency," which leads some candidates to choose Real-time Inference. However, "near real-time" in the context of Asynchronous Inference means the system can handle the request with reasonable latency given the constraints, not that it provides immediate responses like Real-time Inference. Real-time Inference simply cannot handle 1GB payloads or 1-hour processing times effectively.
Official Reference
Exam Strategy
When you see specific numerical constraints in AWS questions (like '1GB payload' and '1 hour processing'), match these exactly to the service capabilities documented by AWS. Don't let qualitative terms like 'near real-time' override the quantitative requirements - Asynchronous Inference is explicitly designed for this exact combination of large payloads, long processing times, and near real-time latency.
Related Analysis
Practice All AIF-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full AIF-C01 Practice Test →