Managed content moderation is available on request. Contact AI Reserve to
define your organization’s policy and have it enabled — nothing changes for
your traffic until you do.
guardrailConfig in a
Converse request is rejected honestly rather than silently ignored). Managed
moderation is the supported path: AI Reserve provisions and manages the
policy in its own infrastructure and applies it to your traffic at the
gateway — so on the covered endpoints, one policy applies to every model you
route through them, not just Bedrock-served ones.
How a block looks
On/v1/chat/completions, blocked requests answer 400 with a stable error
code your integration can branch on:
categories names the policy classes that fired (for example VIOLENCE,
PROMPT_ATTACK, or topic:<your denied topic>) — never the content itself.
On the Bedrock Converse surface, blocks arrive in the AWS error dialect your
SDK already understands: a 400 ValidationException whose message is the
same block copy shown above. The CONTENT_BLOCKED_BY_POLICY code and
categories detail are not carried in the AWS wire shape — if your
integration needs to branch on them, use /v1/chat/completions.
Behavior details
- Off by default. Moderation runs only for organizations that requested it. Everyone else’s requests are untouched — no added latency, no policy calls.
- Input screening, before dispatch. The user-authored text of a request is screened before any model call, wallet charge, or content logging — blocked input is never sent to a provider or persisted. Streaming requests are blocked before the stream opens; served streams are never interrupted.
- Coverage.
/v1/chat/completionsand the Bedrock Converse surface (streaming and non-streaming). The Anthropic-native (/v1/messages) and Responses (/v1/responses) endpoints are not yet screened — if your traffic uses them, tell us when you request enablement and we’ll scope them in. - Fails open, loudly. If the moderation backend is unreachable, requests serve unmoderated rather than failing — availability first — and each degradation window is recorded in the platform’s audit log.
- Content-free audit trail. Blocks are logged as policy category + organization + request id. The screened text itself is never written to moderation logs.
Set it up
- Contact AI Reserve (your account contact, or the address on the portal’s Help page) and describe the policy you need — denied topics, content filters, and how strict each should be.
- We provision the policy, enable it for your organization, and confirm.
- Verify with a request that violates the policy — it should answer
400 CONTENT_BLOCKED_BY_POLICY.