OpenAI Moderation
Hosted moderation API flagging content across 13 harm categories with calibrated scores.
Wraps OpenAI's hosted moderation classifier (default model: omni-moderation-latest) to flag content across 13 harm sub-categories: hate, hate/threatening, harassment, harassment/threatening, self-harm, self-harm/intent, self-harm/instructions, sexual, sexual/minors, violence, violence/graphic, illicit, and illicit/violent. The classifier returns a calibrated per-category probability score alongside a boolean flagged verdict.
Expected input: a single string (or a list of strings, screened one at a time via the batched ThreeStageGuardrail.validate).
GuardrailOutput mapping: - valid is False when OpenAI flags the content or when the maximum per-category score exceeds threshold (otherwise True). - score is the maximum per-category probability (higher = riskier). - categories is the full per-category breakdown: each CategoryResult carries the calibrated score and a triggered flag (set when OpenAI flagged that category or its score exceeds threshold). - raw is the full OpenAI SDK moderation response object.
The current omni-moderation model is a GPT-4o-derived multimodal classifier; the original methodology (taxonomy, active-learning loop, calibration) is described in Markov et al. 2023 (AAAI 2023). The Moderation API is free and does not count toward standard usage quotas.
For more information, see:
Supported Models
omni-moderation-latestomni-moderation-2024-09-26text-moderation-latest
Constructor
model_id
`str
None`
No
None
api_key
`str
None`
No
None
base_url
`str
None`
No
None
threshold
float
No
0.5
Maximum per-category score above which content is considered invalid even if OpenAI did not explicitly flag it. Defaults to 0.5.
Initialize the OpenAI Moderation guardrail.
validate
Default validation pipeline: preprocess -> inference -> postprocess.
Parameters
input_text
`str
list[str]`
Yes
—
Returns: GuardrailOutput | list[GuardrailOutput]
Benchmarks
No benchmark results recorded yet. See the benchmark methodology for how numbers are harvested (published) or measured and added.
License
Vendor: OpenAI
Default license:
proprietary(of the default model/service)
Last updated