What drives inference costs in Amazon Bedrock LLMs?

Generative AI Fundamentals

A company wants to assess the costs that are associated with using a large language model (LLM) to generate inferences. The company wants to use Amazon Bedrock to build generative AI applications. Which factor will drive the inference costs?

  1. Number of tokens consumed Source Reference Answer
  2. Temperature value
  3. Amount of data used to train the LLM
  4. Total training time

Community Votes

A
100%

100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests the distinction between inference pricing models and training cost factors, highlighting that users pay for what they process (tokens), not how the model was built.

Inference costs in Amazon Bedrock are primarily driven by the number of tokens consumed during input and output processing. The community consensus confirms that token volume, not training metrics or model parameters like temperature, determines pricing.

Learners often confuse Option C (Amount of data used to train) or Option D (Total training time) with operational costs, failing to recognize that these are sunk costs associated with model development rather than per-request inference fees.

Community Discussion (4 comments)

Jessiii 👍 1 Selected: A
In the context of using Amazon Bedrock and generative AI models, inference costs are typically driven by the number of tokens consumed during the input and output processing. Number of tokens consumed refers to how many tokens (words, subwords, characters) the model processes during inference (both input and output). More tokens mean higher processing and hence higher costs.
OnePG 👍 2 Selected: A
A. Number of tokens consumed. More tokens used = higher cost. All other affects training costs, not inference costs. Correct answer is A
85b5b55 👍 1 Selected: A
No. of tokens consumed while processing. Tokens are the basic units of input and output that a generative AI model operates on, representing words, subwords, or other linguistic units.
PHD_CHENG 👍 3 Selected: A
A is correct. Token is the basic unit of generative AI model

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option A is correct because Amazon Bedrock, like most generative AI services, charges based on the volume of tokens processed. Tokens represent the basic units of text (words, subwords, or characters) that the model reads as input and generates as output. Since computational resources are directly proportional to the amount of data processed, the number of tokens is the primary driver of inference costs.

Why the Other Options Are Wrong

Options C and D relate to the pre-training phase, which involves building the model before it is available for public use; companies do not pay per-inference for the historical training data or time. Option B (Temperature) is a hyperparameter that controls the randomness of the output but does not affect the computational load or the number of tokens generated, thus having no impact on cost.

Community Comment Notes

Commenters unanimously agree that tokens are the fundamental unit of billing in generative AI. Several comments emphasize that while training costs exist, they are separate from the operational expenditure (OpEx) of running inferences. The consensus reinforces that input and output token counts are the sole metrics for calculating inference bills.

Related Analysis

Practice All AIF-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full AIF-C01 Practice Test →

← Back to AIF-C01 Study Guide