Which metric evaluates fine-tuned LLM accuracy for help desk Q&A?
A company has fine-tuned a large language model (LLM) to answer questions for a help desk. The company wants to determine if the fine-tuning has enhanced the model's accuracy. Which metric should the company use for the evaluation?
Community Votes
100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
This question tests the ability to choose an evaluation metric that balances both precision (avoiding irrelevant answers) and recall (not missing correct answers), which is critical in help desk scenarios.
The F1 score is the recommended metric for evaluating a fine-tuned LLM's accuracy on help desk question-answering tasks because it balances precision and recall. Community candidates unanimously agree that F1 provides the most comprehensive assessment of both correctness and completeness.
Candidates often choose 'Precision' because it directly relates to accuracy, but precision alone ignores recall, meaning a model could give few but correct answers while missing many valid responses.
Community Discussion (5 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Understanding the Evaluation Challenge
When a company fine-tunes a Large Language Model (LLM) for help desk question answering, the goal is to ensure the model provides accurate, relevant, and complete responses. Evaluation metrics must capture both dimensions of quality.
Why F1 Score is Correct
The F1 score is the harmonic mean of precision and recall, providing a single balanced metric that penalizes models that are either:
- Overly conservative (high precision but low recall — misses many correct answers)
- Overly generous (high recall but low precision — includes irrelevant or incorrect information)
Why Other Options Fail
- A. Precision: Measures only the proportion of correct answers among all answers given. A model could achieve high precision by answering very few questions, ignoring many valid queries.
- B. Time to first token: This is a latency/performance metric, not an accuracy metric. It measures speed, not correctness.
- D. Word error rate (WER): Primarily used in speech recognition to measure transcription accuracy by comparing word sequences. It is not suitable for evaluating semantic correctness in Q&A tasks.
Community Consensus
All community voters selected C. F1 score, emphasizing that it provides the most balanced evaluation for question-answering tasks where both precision and recall are critical.
Official Reference
Exam Strategy
When a question asks about evaluating model 'accuracy' in a balanced context (e.g., Q&A, classification), always consider metrics that combine multiple dimensions like F1 score. Eliminate latency metrics (like time to first token) immediately when the question focuses on accuracy or correctness.
Related Analysis
Practice All AIF-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full AIF-C01 Practice Test →