gpt-oss-safeguard-20b

Open safety-reasoning model that applies your own written policy.

gpt-oss-safeguard-20b is an open-weight safety reasoning model from OpenAI built on gpt-oss-20b. Instead of a fixed taxonomy it reasons over a policy you supply at inference time, for content classification, LLM filtering and trust & safety labelling.

Provider
OpenAI
Type
Safety classifier
Released
Oct 29, 2025
Context window
131K tokens
Max output
66K tokens
Price
$0.075 input / $0.30 output per 1M tokens
Input
text
Output
text
Open weights
Yes
License
Apache 2.0

Best for

Strengths

More from OpenAI

Sources