Block Jailbreak Attacks in Azure AI Content Safety
You have an Azure subscription that contains an Azure OpenAI resource named AI1. You build a chatbot that uses AI1 to provide generative answers to specific questions. You need to ensure that questions intended to circumvent built-in safety features are blocked. Which Azure AI Content Safety feature should you implement?
Community Votes
100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests identifying the correct Azure AI Content Safety feature for a specific threat type; the common trap is confusing general content moderation with specific jailbreak circumvention.
To prevent users from bypassing built-in safety features in an Azure OpenAI chatbot, you must implement Jailbreak risk detection. This Azure AI Content Safety feature identifies and blocks adversarial prompts designed to circumvent model restrictions.
Choosing Moderate text content (C) because it sounds like a general safety feature, but it fails to specifically target adversarial prompt attacks designed to bypass system constraints.
Community Discussion (3 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Jailbreak risk detection (now known as Prompt Shields) is explicitly designed to identify and block user attempts to manipulate or circumvent the built-in safety mechanisms of an AI model. When a chatbot uses Azure OpenAI, users might try to bypass content restrictions using adversarial prompts, making this feature essential for maintaining system integrity.Why the& Other Options Are Wrong
Monitor online activity (A) is a general monitoring concept and not a specific Azure AI Content Safety feature for blocking adversarial prompts. Moderate text content (C) checks for offensive or inappropriate language but does not detect structured attempts to break out of system instructions. Protected material text detection (D) identifies known copyrighted text but does not prevent safety circumvention.Community Comment Notes
Commenters highlighted that Jailbreak risk detection has been updated and is now referred to as "Prompt shields for user Prompts", as mbsff noted. Another commenter emphasized that this feature is "specifically designed to detect and block user attempts to manipulate or circumvent the built-in safety mechanisms".Official Reference
Exam Strategy
Map the specific threat described in the question to the exact Content Safety feature. Circumventing safety features always points to jailbreak detection, while offensive language points to moderation.