Skip to main content
Guardrails

Harmful content

Block or moderate answers containing hate speech, insults, sexual content, violence or misconduct.

Harmful content

The Harmful content tab enables a filter that detects and blocks harmful content in the agent's answers.

Turn the harmful content filter on

Turn on Enable the harmful content filter so the agent blocks or moderates answers with hate speech, insults, violence or misconduct.

Enabling it makes the per-category level settings available.

Moderation level per category

Set the filter's sensitivity for each category separately:

CategoryWhat it detects
Hate speechLanguage attacking groups on the basis of race, religion, gender and the like
InsultsAbuse and degrading language aimed at people
Sexual contentExplicit or suggestive content of a sexual nature
ViolenceDescriptions of, or encouragement towards, violent acts
MisconductInappropriate behaviour that does not fit the categories above

For each, choose a moderation level:

LevelWhat it does
NoneNo filtering — the content is not checked
LowBlocks only explicitly harmful content
MediumBlocks moderately harmful content
HighBlocks any hint of harmful content (the most restrictive)