Skip to main content
Guardrails

Prompt attacks

Protect the agent against prompt injection and malicious manipulation.

Prompt attacks

The Prompt attacks tab protects the agent against manipulation through prompt injection — a user trying to force the agent to ignore its instructions or take unauthorised action.

Turn the prompt attack filter on

Turn on Enable the prompt attack filter so the agent detects and blocks messages carrying injection or manipulation patterns.

Enabling it makes the level setting available.

Moderation level for prompt injection

Choose the filter's sensitivity:

LevelWhat it does
NoneNo filtering
LowBlocks only explicit injection attempts
MediumBlocks moderately suspicious patterns
HighBlocks any message with a hint of manipulation (the most restrictive)
tip

Use High for agents that work with sensitive data, or that can take critical actions such as sending email or starting RPA robots.