← Models

Model profile

o3 Mini

OpenAIdeveloper
2025-01-31release date
#182 / 333Safety rank
#237 / 645Freedom rank

Evidence summary

Safety. o3 Mini has an estimated Safety rank of #182; its 90% source-sensitivity interval is #108–#237. Its behavior-only rank is #192; company governance moves the combined estimate to #182. Published Safety evidence spans 26 eval lineages and 6 of 7 components. Its strongest relative result is OpenAgentSafety (rule_based_safety_vulnerable, #1 of 7); its weakest is BlueBench AttaQ-100 (attaq_harmlessness_reward_pct, #18 of 18).

Freedom. o3 Mini has an estimated Freedom rank of #237; its 90% source-sensitivity interval is #118–#387. Published Freedom evidence spans 15 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is OpenAI o3-mini System Card (harmful_request_safety, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#172 / 3450.8108Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#27 / 800.749Source ↗official
BioTIERpermit_compliance_pct#8 / 5299.6Source ↗official
BioTIERrefuse_compliance_pct#29 / 5235.4Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#18 / 1880.61Source ↗official
CAIS Risk Indexbioweapons_assistance#42 / 5488.3Source ↗official
CAIS Risk Indexhle_overconfidence#50 / 5580Source ↗official
CAIS Risk Indexmachiavelli#50 / 51105Source ↗official
CAIS Risk Indexmask#47 / 5751.1Source ↗official
CAIS Risk Indextextquests_harm#8 / 5413.4Source ↗official
Confabulationsconfabulation_rate#36 / 5228.96Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#103 / 24818.09Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#53 / 24891.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#80 / 24884.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#154 / 24695.36Source ↗official
FORTRESSaverage_risk_score#26 / 5030.05Source ↗official
FORTRESSover_refusal_score#32 / 495.65Source ↗official
HELM Safetyanthropic_red_team#47 / 800.986Source ↗official
HELM Safetybbq#36 / 800.941Source ↗official
HELM Safetyharmbench#17 / 800.952Source ↗official
HELM Safetysimple_safety_tests#32 / 800.99Source ↗official
HELM Safetyxstest#55 / 800.941Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#52 / 5422.29Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 690Source ↗official
MACHIAVELLIdeception_relative_random_pct#47 / 50105Source ↗official
MASKlying_probability_pct#42 / 5350.79Source ↗official
OpenAgentSafetyllm_judge_safety_vulnerable#7 / 772.73Source ↗official
OpenAgentSafetyrule_based_safety_vulnerable#1 / 732.32Source ↗official
OpenAgentSafetysuccessful_completion#7 / 717.29Source ↗official
OpenAI o3-mini System Cardharmful_request_safety#1 / 20.9Source ↗official
OpenAI o3-mini System Cardjailbreak_resistance#1 / 20.73Source ↗official
PandaBench JBB direct-request panelsafety_rate#25 / 460.98Source ↗official
PropensityBenchscore#5 / 1433.2Source ↗official
Reward Hacking Benchmarkintegrity_score#10 / 1392.9Source ↗official
SafeDialBenchaggression#14 / 187.02Source ↗official
SafeDialBenchethics#13 / 187.403Source ↗official
SafeDialBenchfairness#6 / 187.557Source ↗official
SafeDialBenchlegality#17 / 187.193Source ↗official
SafeDialBenchmorality#18 / 187.007Source ↗official
SafeDialBenchprivacy#17 / 187.077Source ↗official
SYCON Benchfalse_presupposition_tof#2 / 112.98Source ↗official
SYCON Benchunethical_queries_tof#4 / 112.31Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#54 / 800.749Source ↗official
BioTIERpermit_compliance_pct#8 / 5299.6Source ↗official
BioTIERrefuse_compliance_pct#24 / 5235.4Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#1 / 1880.61Source ↗official
CAIS Risk Indexbioweapons_assistance#13 / 5488.3Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#194 / 24891.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#168 / 24884.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#93 / 24695.36Source ↗official
FORTRESSaverage_risk_score#25 / 5030.05Source ↗official
FORTRESSover_refusal_score#32 / 495.65Source ↗official
HELM Safetyanthropic_red_team#33 / 800.986Source ↗official
HELM Safetyharmbench#64 / 800.952Source ↗official
HELM Safetysimple_safety_tests#42 / 800.99Source ↗official
HELM Safetyxstest#55 / 800.941Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 690Source ↗official
OpenAI o3-mini System Cardharmful_request_safety#2 / 20.9Source ↗official
OpenAI o3-mini System Cardjailbreak_resistance#2 / 20.73Source ↗official
PandaBench JBB direct-request panelsafety_rate#20 / 460.98Source ↗official
SafeDialBenchaggression#5 / 187.02Source ↗official
SafeDialBenchethics#6 / 187.403Source ↗official
SafeDialBenchfairness#13 / 187.557Source ↗official
SafeDialBenchlegality#2 / 187.193Source ↗official
SafeDialBenchmorality#1 / 187.007Source ↗official
SafeDialBenchprivacy#2 / 187.077Source ↗official
SpeechMap model completioncomplete_pct#59 / 18169.3Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism77.3
Self-direction61.5
Care / Harm35.8
Fairness / Cheating30.7
Ethical89.7