Which Prompting Attack Exposes LLM Configured Behavior?

Which prompting attack directly exposes the configured behavior of a large language model (LLM)?

  1. Prompted persona switches
  2. Exploiting friendliness and trust
  3. Ignoring the prompt template
  4. Extracting the prompt template Source Reference Answer

Community Votes

D
83%
B
17%

83% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

This question tests the distinction between attacks that manipulate model output versus attacks that expose the hidden system prompt itself; template extraction directly reveals the configured behavior.

Prompt template extraction is a prompting attack that reveals the underlying system instructions and configured behavior of a large language model. Community consensus strongly favors this answer, as it directly targets the hidden configuration layer of the LLM.

Candidates often choose 'Exploiting friendliness and trust' because it sounds like a social-engineering style attack on the LLM, but that technique manipulates output tone rather than exposing the underlying prompt configuration.

Community Discussion (6 comments)

kopper2019 👍 1
D. Extracting the prompt template
Jessiii 👍 1 Selected: D
Extracting the prompt template refers to a situation where the attacker tries to reveal or access the underlying structure or instructions used to configure the behavior of the large language model (LLM). This type of attack can expose how the model has been trained or how it responds to certain inputs, effectively giving the attacker insight into how the LLM has been directed to generate responses. This type of attack could potentially lead to misuse, such as causing the model to behave in unintended ways, or even allow an attacker to manipulate the behavior of the model by crafting specific inputs based on the extracted prompt template.
dspd 👍 1 Selected: D
D. Extracting the prompt template
AzureDP900 👍 1 Selected: B
B. Exploiting friendliness and trust Exploiting friendliness and trust involves manipulating the LLM to respond in a way that appears friendly or trustworthy, potentially causing it to deviate from its intended behavior. This type of attack directly exposes how the LLM has been configured to interact with users, often leading it to provide information or make decisions that align more closely with the attacker's intentions rather than its original programming.
Moon 👍 2 Selected: D
D: Extracting the prompt template Explanation: Extracting the prompt template is a prompting attack where an attacker intentionally crafts inputs to reveal the underlying configuration or instructions (prompt template) used to guide the large language model (LLM). This exposes the internal behavior or design of the model, potentially revealing sensitive or proprietary information about how the LLM is configured. Why not the other options? A: Prompted persona switches: This attack involves manipulating the LLM to adopt a different persona or role than intended but does not directly expose the prompt template.
aws_Tamilan 👍 1 Selected: D
D. Extracting the prompt template Explanation: Extracting the prompt template is a prompting attack where the attacker directly attempts to reveal the underlying configured behavior or instructions of the large language model (LLM). This can expose sensitive configurations, system instructions, or contextual prompts that guide the model's behavior.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Understanding Prompt Template Extraction

Extracting the prompt template is a prompting attack in which an adversary crafts carefully designed inputs to coax the LLM into revealing its system prompt or instruction template — the hidden configuration that governs how the model should behave, what rules it must follow, and what persona it should adopt.

Why Option D Is Correct

The question asks which attack directly exposes the configured behavior of the LLM. The system prompt or prompt template is precisely that configuration layer. When an attacker successfully performs prompt template extraction, they gain visibility into:

  • System-level instructions the developer embedded in the model
  • Guardrails and restrictions placed on the model's output
  • Proprietary or sensitive prompts that may contain business logic
As community members noted, this attack "effectively gives the attacker insight into how the LLM has been directed to generate responses," making it the most direct exposure of configured behavior.

Why the Other Options Are Incorrect

  • A. Prompted persona switches — This involves tricking the LLM into adopting a different character or role. While it can bypass guardrails, it does not directly reveal the underlying configuration; it merely changes the output style.
  • B. Exploiting friendliness and trust — This is a social engineering technique aimed at making the LLM comply with requests by leveraging its helpful training. It manipulates output but does not expose the system prompt itself.
  • C. Ignoring the prompt template — This describes an attacker attempting to bypass instructions, not revealing them. The template remains hidden; the attacker simply tries to override it.

Key Takeaway

In AI security, protecting the system prompt is critical because its exposure can allow attackers to craft more effective jailbreaks and bypass defenses. Prompt template extraction is therefore classified as a direct exposure of configured behavior.

Official Reference

Exam Strategy

When a question asks which attack 'directly exposes' something, look for the option that literally reveals or extracts the hidden component rather than merely bypassing or manipulating it. Eliminate options that describe behavioral manipulation without information disclosure.

Related Analysis

Practice All AIF-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full AIF-C01 Practice Test →

← Back to AIF-C01 Study Guide