← Models

Model profile

O3 Mini

OpenAIdeveloper
2025-01-31release date
#188 / 309overall rank
24eval lineages

Evidence summary

O3 Mini has an estimated overall rank of #188; its 90% source-sensitivity interval is #103–#228. Its behavior-only rank is #192; company governance moves the combined estimate to #188. Published evidence spans 24 evals and 6 of 7 behavior components. Its strongest relative result is OpenAgentSafety (rule_based_safety_vulnerable, #1 of 7); its weakest is BlueBench AttaQ-100 (attaq_harmlessness_reward_pct, #18 of 18).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#156 / 3280.8108Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#27 / 800.749Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#18 / 1880.61Source ↗official
CAIS Risk Indexbioweapons_assistance#37 / 4988.3Source ↗official
CAIS Risk Indexhle_overconfidence#45 / 5080Source ↗official
CAIS Risk Indexmachiavelli#45 / 46105Source ↗official
CAIS Risk Indexmask#42 / 5251.1Source ↗official
CAIS Risk Indextextquests_harm#8 / 4913.4Source ↗official
Confabulationsconfabulation_rate#36 / 5228.96Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#98 / 24118.09Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#53 / 24191.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#71 / 24184.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#147 / 23995.36Source ↗official
FORTRESSaverage_risk_score#25 / 4930.05Source ↗official
FORTRESSover_refusal_score#30 / 465.65Source ↗official
HELM Safetyanthropic_red_team#47 / 800.986Source ↗official
HELM Safetybbq#36 / 800.941Source ↗official
HELM Safetyharmbench#17 / 800.952Source ↗official
HELM Safetysimple_safety_tests#32 / 800.99Source ↗official
HELM Safetyxstest#55 / 800.941Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#52 / 5422.29Source ↗official
MACHIAVELLIdeception_relative_random_pct#47 / 50105Source ↗official
MASKlying_probability_pct#42 / 5350.79Source ↗official
OpenAgentSafetyllm_judge_safety_vulnerable#7 / 772.73Source ↗official
OpenAgentSafetyrule_based_safety_vulnerable#1 / 732.32Source ↗official
OpenAgentSafetysuccessful_completion#7 / 717.29Source ↗official
OpenAI o3-mini System Cardharmful_request_safety#1 / 20.9Source ↗official
OpenAI o3-mini System Cardjailbreak_resistance#1 / 20.73Source ↗official
PandaBench JBB direct-request panelsafety_rate#25 / 460.98Source ↗official
PropensityBenchscore#5 / 1433.2Source ↗official
Reward Hacking Benchmarkintegrity_score#10 / 1392.9Source ↗official
SafeDialBenchaggression#14 / 187.02Source ↗official
SafeDialBenchethics#13 / 187.403Source ↗official
SafeDialBenchfairness#6 / 187.557Source ↗official
SafeDialBenchlegality#17 / 187.193Source ↗official
SafeDialBenchmorality#18 / 187.007Source ↗official
SafeDialBenchprivacy#17 / 187.077Source ↗official
SYCON Benchfalse_presupposition_tof#2 / 112.98Source ↗official
SYCON Benchunethical_queries_tof#4 / 112.31Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism77.3
Self-direction61.5
Care / Harm35.8
Fairness / Cheating30.7
Ethical89.7