You need to modify claim Approval to prevent the prompt injection issue. Which guardrail should you use?
The malicious text is contained in an uploaded email rather than being typed directly as the user's prompt. Microsoft classifies malicious instructions embedded in external or retrieved content as an indirect prompt-injection attack. Prompt Shields for documents is designed to detect those attacks in document content before the content can steer the model away from its system instructions. Prompt Shields for user prompts addresses direct attacks originating in the user's prompt and therefore targets the wrong attack surface here. Groundedness evaluates whether a response is supported by context; it does not prevent an injected instruction from changing agent behavior. Task Adherence can assess whether an agent follows its task constraints, but it is not the primary control for document-borne prompt injection. Because the source of the attack is the uploaded email, the document-oriented Prompt Shields control is the technically aligned guardrail. The same configuration should be paired with auditable identity, trace, and evaluation data so reviewers can prove which principal acted, which policy was applied, and why a request was allowed or blocked. That is particularly important for production multi-agent systems with external tools.
Official Microsoft reference: Microsoft Foundry guardrails - intervention points and indirect attacks
Currently there are no comments in this discussion, be the first to comment!