Prompt attacks
Protect the agent against prompt injection and malicious manipulation.
Prompt attacks
The Prompt attacks tab protects the agent against manipulation through prompt injection — a user trying to force the agent to ignore its instructions or take unauthorised action.
Turn the prompt attack filter on
Turn on Enable the prompt attack filter so the agent detects and blocks messages carrying injection or manipulation patterns.
Enabling it makes the level setting available.
Moderation level for prompt injection
Choose the filter's sensitivity:
| Level | What it does |
|---|---|
| None | No filtering |
| Low | Blocks only explicit injection attempts |
| Medium | Blocks moderately suspicious patterns |
| High | Blocks any message with a hint of manipulation (the most restrictive) |
tip
Use High for agents that work with sensitive data, or that can take critical actions such as sending email or starting RPA robots.