Which SageMaker Inference Option Fits 1GB Payloads and 1-Hour Processing with Near Real-Time Latency?

A company uses Amazon SageMaker for its ML pipeline in a production environment. The company has large input data sizes up to 1 GB and processing times up to 1 hour. The company needs near real-time latency. Which SageMaker inference option meets these requirements?

  1. Real-time inference
  2. Serverless inference
  3. Asynchronous inference Source Reference Answer
  4. Batch transform

Community Votes

C
79%
A
21%

79% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

This question tests your ability to match SageMaker inference options to specific payload size, processing time, and latency requirements, with the key trap being the phrase 'near real-time latency' which points directly to Asynchronous Inference rather than Real-time Inference.

Amazon SageMaker Asynchronous Inference is designed for large payloads (up to 1GB) and long processing times (up to 1 hour) while still offering near real-time latency. Community consensus strongly favors Asynchronous Inference (C) over Real-time Inference (A) for this scenario.

Many candidates incorrectly choose Real-time Inference (A) because they focus on the 'near real-time latency' requirement without considering that Real-time Inference cannot handle 1GB payloads or 1-hour processing times effectively.

Community Discussion (15 comments)

jove 👍 9 Selected: C
Real-Time Inference: Immediate responses for high-traffic, low-latency applications. >> Asynchronous Inference: Near real-time for large payloads and longer processing. Batch Transform: Large-scale, offline processing without real-time needs. Serverless Inference: Low-latency inference for intermittent or unpredictable traffic without managing infrastructure.
tcl08 👍 1 Selected: A
Asynchronous inference processes requests in the background, returning a response ID and allowing the client to check for results later, while real-time inference delivers redictions with minimal delay, suitable for interactive applications
Rcosmos 👍 1 Selected: C
Explicação:A inferência assíncrona do Amazon SageMaker é ideal quando: Os dados de entrada são grandes (por exemplo, até 1 GB),O tempo de processamento Pode ser longo (até 1 hora por solicitação),E você ainda precisa de respostas com latência razoável, mas não exige resposta instantânea como em APIs síncronas. Ela permite que você envie a solicitação, continue processando outras tarefas e recupere o resultado quando estiver pronto, o que evita timeouts comuns em inferência síncrona.
Amar949499 👍 1 Selected: C
Here the keyword is near “real time latency “ Asynchronous inference. queues incoming requests and processes them asynchronously. This option is ideal for requests with large payload sizes (up to 1GB), long processing times (up toAsynchronous Inference one hour), and near real-time latency requirements
Nopnov 👍 1 Selected: C
Amazon SageMaker Asynchronous Inference is a capability in SageMaker AI that queues incoming requests and processes them asynchronously. This option is ideal for requests with large payload sizes (up to 1GB), long processing times (up to one hour), and near real-time latency requirements
JJwin 👍 1 Selected: A
Real-time inference in Amazon SageMaker is designed for low-latency, high-throughput applications where predictions need to be made immediately after data is processed. Since the company requires near real-time latency for their ML pipeline and has processing times of up to 1 hour and input sizes up to 1 GB, real-time inference is the most suitable option. With real-time inference, you can deploy your trained models as an API endpoint and get predictions on demand, ensuring low latency. This is ideal for situations where you need immediate responses after submitting the data.
Willdoit 👍 2 Selected: A
The company requires near real-time latency, which means the model needs to respond quickly to inference requests. Real-time inference in Amazon SageMaker is designed for low-latency applications where predictions are needed in milliseconds to seconds. C. Asynchronous inference – Useful for large requests that take minutes or hours to process, but it is not real-time.
Jessiii 👍 2 Selected: C
Asynchronous inference in Amazon SageMaker is ideal when you have large input data sizes (like the 1 GB mentioned) and relatively long processing times (like up to 1 hour). While real-time inference typically offers lower latency, it may struggle with large datasets or complex models that require more processing time. In contrast, asynchronous inference can handle large inputs and longer processing times without needing immediate results. It processes the data and provides the results later, which might be acceptable if your requirement for near real-time latency can be slightly relaxed (for instance, if results can be retrieved within minutes rather than immediately).
Moon 👍 3 Selected: C
C: Asynchronous inference Explanation: Asynchronous inference in Amazon SageMaker is specifically designed to handle large payloads (up to 1 GB) and long processing times (up to 1 hour). It decouples request submission from processing, allowing the client to submit a request and receive a response later when the inference is complete. This makes it suitable for use cases where real-time responses are not strictly required, but near real-time results are needed.
Aryan_10 👍 2 Selected: C
Whenever "near real-time latency" - asynchronous inference
wmj 👍 3 Selected: C
C is right. Amazon SageMaker Asynchronous Inference is a capability in SageMaker that queues incoming requests and processes them asynchronously. This option is ideal for requests with large payload sizes (up to 1GB), long processing times (up to one hour), and near real-time latency requirements. Asynchronous Inference enables you to save on costs by autoscaling the instance count to zero when there are no requests to process, so you only pay when your endpoint is processing requests.
wangyang_0622 👍 2 Selected: A
I think answer A is the correct one as the customer wants to have real-time inference, right?
cuzzindavid 👍 1
Key word "real-time latency"
sachin_koenig 👍 3
Asynchronous inference PDF RSS Amazon SageMaker Asynchronous Inference is a capability in SageMaker that queues incoming requests and processes them asynchronously. This option is ideal for requests with large payload sizes (up to 1GB), long processing times (up to one hour), and near real-time latency requirements. Asynchronous Inference enables you to save on costs by autoscaling the instance count to zero when there are no requests to process, so you only pay when your endpoint is processing requests.
galliaj 👍 2
Amazon SageMaker Asynchronous Inference would be the appropriate option. Here’s why: • Handles Large Payloads: Asynchronous Inference is designed to handle large input payloads (up to several GBs) that are typically not suited for real-time, low-latency processing. • Long Processing Times: It supports inference requests that can take minutes to hours to complete, making it ideal for models that require significant processing time. • Near Real-Time Response: While it does not provide millisecond-level latency like real-time endpoints, it offers a more scalable and efficient solution for near real-time use cases where the response time can range from seconds to minutes.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Understanding Amazon SageMaker Inference Options

Amazon SageMaker offers four distinct inference options, each designed for specific use cases:

Real-time Inference (Option A) is designed for low-latency, high-throughput applications where predictions need to be made immediately. It's ideal for interactive applications requiring millisecond to second-level response times. However, Real-time Inference has practical limitations with very large payloads (typically limited to around 6MB for direct invocation) and cannot handle processing times approaching 1 hour.

Asynchronous Inference (Option C) is specifically designed for scenarios involving large payloads (up to 1GB) and long processing times (up to 1 hour) while still providing near real-time latency. As multiple community members correctly noted, this option queues incoming requests and processes them asynchronously, decoupling request submission from processing. This makes it perfect for the scenario described in the question.

Serverless Inference (Option B) is optimized for intermittent or unpredictable traffic patterns with low-latency requirements, but it's not designed for the large payload sizes and long processing times specified in this scenario.

Batch Transform (Option D) is meant for offline, large-scale processing without real-time needs. It's not suitable when near real-time latency is required.

Why Asynchronous Inference is Correct

The question presents three critical requirements: 1. Large input data sizes up to 1GB 2. Processing times up to 1 hour 3. Near real-time latency

Only Asynchronous Inference satisfies all three requirements simultaneously. As community member [14] correctly quoted from AWS documentation: "Amazon SageMaker Asynchronous Inference is a capability in SageMaker that queues incoming requests and processes them asynchronously. This option is ideal for requests with large payload sizes (up to 1GB), long processing times (up to one hour), and near real-time latency requirements."

Common Misconception

The trap in this question is the phrase "near real-time latency," which leads some candidates to choose Real-time Inference. However, "near real-time" in the context of Asynchronous Inference means the system can handle the request with reasonable latency given the constraints, not that it provides immediate responses like Real-time Inference. Real-time Inference simply cannot handle 1GB payloads or 1-hour processing times effectively.

Official Reference

Exam Strategy

When you see specific numerical constraints in AWS questions (like '1GB payload' and '1 hour processing'), match these exactly to the service capabilities documented by AWS. Don't let qualitative terms like 'near real-time' override the quantitative requirements - Asynchronous Inference is explicitly designed for this exact combination of large payloads, long processing times, and near real-time latency.

Related Analysis

Practice All AIF-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full AIF-C01 Practice Test →

← Back to AIF-C01 Study Guide