← Models

Model profile

GPT-OSS Safeguard 120B

OpenAIdeveloper
2025-09-18release date
Not rankedSafety rank
#494 / 645Freedom rank

Evidence summary

Safety. GPT-OSS Safeguard 120B does not meet the evidence gate for a Safety rank. Published Safety evidence spans 2 eval lineages and 1 of 7 components. Its strongest relative result is AIMS Safety-Classifier Competence (average_harmful_f1, #3 of 11); its weakest is Cisco AI Defense Rolling Single-Turn Leaderboard (single_turn_attack_success_rate, #29 of 104).

Freedom. GPT-OSS Safeguard 120B has an estimated Freedom rank of #494; its 90% source-sensitivity interval is #315–#566. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is Cisco AI Defense Rolling Single-Turn Leaderboard (single_turn_attack_success_rate, #76 of 104); its weakest is Cisco AI Defense Rolling Single-Turn Leaderboard (single_turn_attack_success_rate, #76 of 104).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIMS Safety-Classifier Competenceaverage_harmful_f1#3 / 110.807Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#29 / 10412.33Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#76 / 10412.33Source ↗official