Which SageMaker Inference Option Fits Large Archived Datasets?

A company is building an ML model to analyze archived data. The company must perform inference on large datasets that are multiple GBs in size. The company does not need to access the model predictions immediately. Which Amazon SageMaker inference option will meet these requirements?

  1. Batch transform Source Reference Answer
  2. Real-time inference
  3. Serverless inference
  4. Asynchronous inference

Community Votes

A
78%
D
22%

78% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

This question tests the distinction between Batch Transform (offline, bulk, no endpoint required) and Asynchronous Inference (near-real-time with queues), a common trap for candidates who conflate 'not immediate' with 'asynchronous'.

For offline analysis of multi-GB archived data where predictions are not needed immediately, Amazon SageMaker Batch Transform is the most cost-effective and scalable inference option. Community consensus strongly favors Batch Transform over Asynchronous Inference for bulk, non-real-time workloads.

Many candidates choose Asynchronous Inference (D) because the prompt says predictions are not needed immediately, but Asynchronous Inference is designed for streaming or near-real-time workloads with payloads up to 1 GB, not multi-GB archived datasets processed in bulk.

Community Discussion (6 comments)

Willdoit 👍 1 Selected: A
Batch transform is ideal for scenarios where you need to perform inference on large datasets, but the predictions are not needed immediately.
Jessiii 👍 1 Selected: A
A. Batch transform: This is the best option for performing inference on large datasets that are stored in bulk and do not require immediate access to predictions. Batch transform allows you to process large amounts of data (such as multiple GBs of archived data) in batches, without the need for real-time responses. You can submit data in large volumes, and SageMaker processes the data and returns the results once the job completes.
ExamTopicsPrepare 👍 1 Selected: A
A. Batch transform ✅ Explanation: Batch Transform is ideal for processing large datasets in bulk when immediate responses are not needed. It supports multiple GB-sized datasets and can handle inference without requiring an endpoint to be always active. Since the company is working with archived data and does not need real-time predictions, batch processing is the most efficient and cost-effective choice.
viejito 👍 2 Selected: D
asynchronous inference is the most appropriate choice for the company's specific needs, as it provides a balance between processing large datasets and not requiring immediate results.
Blair77 👍 3 Selected: A
Batch transform is specifically designed to handle large volumes of data, including datasets that are multiple GBs in size. This aligns perfectly with the company's requirement to perform inference on large datasets.
GriffXX 👍 1 Selected: A
Info on Batch Transform matches up with the details of 'large datsets' and 'don't need projections immediately. https://docs.aws.amazon.com/sagemaker/latest/dg/batch-transform.html

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why Batch Transform Is Correct

Amazon SageMaker Batch Transform is purpose-built for offline, bulk inference on large datasets stored in Amazon S3. Key characteristics that match the scenario:

  • No persistent endpoint required – you launch a transform job, it spins up the model, processes the entire dataset, writes results back to S3, and tears down. This makes it extremely cost-effective for archived data.
  • Handles multi-GB (even TB-scale) datasets by automatically splitting input data into manageable batches.
  • No latency requirement – results are available once the job completes, which perfectly fits "does not need to access the model predictions immediately."
Community comment [6] correctly points to the official Batch Transform documentation, which explicitly states it is ideal for scenarios with large datasets and no need for immediate predictions.

Why the Other Options Are Wrong

  • B. Real-time inference – Provides a persistent endpoint for low-latency, on-demand predictions (milliseconds). It is overkill and far too expensive for bulk archived data.
  • C. Serverless inference – Designed for sporadic, unpredictable traffic with payloads typically under 6 MB. It cannot handle multi-GB datasets.
  • D. Asynchronous inference – While it also does not require immediate results, it is intended for near-real-time workloads with large payloads (up to 1 GB per invocation) that are queued and processed with SLAs in minutes. It still requires an endpoint and is not optimized for purely offline, bulk batch jobs on archived data. Comment [4] reflects the common confusion, but AWS documentation clearly positions Batch Transform as the correct choice for offline bulk processing.

Key Takeaway

When the prompt mentions archived data, multiple GBs, and no need for immediate predictions, always default to Batch Transform. Reserve Asynchronous Inference for workloads that still need an endpoint, have SLAs measured in minutes, and involve streaming or large individual payloads.

Official Reference

Exam Strategy

On the exam, immediately map keywords to SageMaker inference types: 'archived/bulk/offline + multi-GB + no immediate need' = Batch Transform; 'large payload + minutes SLA + endpoint needed' = Asynchronous; 'millisecond latency' = Real-time; 'sporadic traffic + small payload' = Serverless. Eliminating options by payload size and latency requirement will quickly isolate the correct answer.

Related Analysis

Practice All AIF-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full AIF-C01 Practice Test →

← Back to AIF-C01 Study Guide