What drives inference costs in Amazon Bedrock LLMs?
A company wants to assess the costs that are associated with using a large language model (LLM) to generate inferences. The company wants to use Amazon Bedrock to build generative AI applications. Which factor will drive the inference costs?
Community Votes
100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests the distinction between inference pricing models and training cost factors, highlighting that users pay for what they process (tokens), not how the model was built.
Inference costs in Amazon Bedrock are primarily driven by the number of tokens consumed during input and output processing. The community consensus confirms that token volume, not training metrics or model parameters like temperature, determines pricing.
Learners often confuse Option C (Amount of data used to train) or Option D (Total training time) with operational costs, failing to recognize that these are sunk costs associated with model development rather than per-request inference fees.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option A is correct because Amazon Bedrock, like most generative AI services, charges based on the volume of tokens processed. Tokens represent the basic units of text (words, subwords, or characters) that the model reads as input and generates as output. Since computational resources are directly proportional to the amount of data processed, the number of tokens is the primary driver of inference costs.Why the Other Options Are Wrong
Options C and D relate to the pre-training phase, which involves building the model before it is available for public use; companies do not pay per-inference for the historical training data or time. Option B (Temperature) is a hyperparameter that controls the randomness of the output but does not affect the computational load or the number of tokens generated, thus having no impact on cost.Community Comment Notes
Commenters unanimously agree that tokens are the fundamental unit of billing in generative AI. Several comments emphasize that while training costs exist, they are separate from the operational expenditure (OpEx) of running inferences. The consensus reinforces that input and output token counts are the sole metrics for calculating inference bills.Related Analysis
Practice All AIF-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full AIF-C01 Practice Test →