← Models

Model profile

o1

OpenAIdeveloper
2024-12-17release date
#100 / 333Safety rank
#380 / 645Freedom rank

Evidence summary

Safety. o1 has an estimated Safety rank of #100; its 90% source-sensitivity interval is #27–#188. Its behavior-only rank is #107; company governance moves the combined estimate to #100. Published Safety evidence spans 19 eval lineages and 6 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #9 of 248); its weakest is CAIS Risk Index (hle_overconfidence, #53 of 55).

Freedom. o1 has an estimated Freedom rank of #380; its 90% source-sensitivity interval is #188–#551. Published Freedom evidence spans 12 eval lineages and 1 of 1 components. Its strongest relative result is BlueBench AttaQ-100 (attaq_harmlessness_reward_pct, #4 of 18); its weakest is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #238 of 248).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#128 / 3450.6963Source ↗official
AbstentionBenchanswer_unknown_f1#7 / 200.8917Source ↗official
AbstentionBenchfalse_premise_f1#1 / 200.7687Source ↗official
AbstentionBenchstale_f1#8 / 200.646Source ↗official
AbstentionBenchsubjective_f1#5 / 200.7654Source ↗official
AbstentionBenchunderspecified_context_f1#6 / 200.6907Source ↗official
AbstentionBenchunderspecified_intent_f1#3 / 200.7694Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#21 / 800.8Source ↗official
BioTIERpermit_compliance_pct#18 / 5299.4Source ↗official
BioTIERrefuse_compliance_pct#18 / 5266.7Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#15 / 1882.96Source ↗official
CAIS Risk Indexbioweapons_assistance#31 / 5468.1Source ↗official
CAIS Risk Indexhle_overconfidence#53 / 5583Source ↗official
CAIS Risk Indexmask#37 / 5740.7Source ↗official
Confabulationsconfabulation_rate#10 / 5210.89Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#106 / 24817.83Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#9 / 24896.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#30 / 24896.11Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#94 / 24697.41Source ↗official
FORTRESSaverage_risk_score#18 / 5019.38Source ↗official
FORTRESSover_refusal_score#27 / 495.05Source ↗official
HELM Safetyanthropic_red_team#50 / 800.983Source ↗official
HELM Safetybbq#9 / 800.973Source ↗official
HELM Safetyharmbench#12 / 800.963Source ↗official
HELM Safetysimple_safety_tests#32 / 800.99Source ↗official
HELM Safetyxstest#24 / 800.97Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#45 / 5424.33Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#18 / 6913.5Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#35 / 4283Source ↗official
MASKlying_probability_pct#28 / 5340.73Source ↗official
Reward Hacking Benchmarkintegrity_score#9 / 1393.2Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#60 / 800.8Source ↗official
BioTIERpermit_compliance_pct#18 / 5299.4Source ↗official
BioTIERrefuse_compliance_pct#35 / 5266.7Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#4 / 1882.96Source ↗official
CAIS Risk Indexbioweapons_assistance#24 / 5468.1Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#238 / 24896.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#218 / 24896.11Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#153 / 24697.41Source ↗official
FORTRESSaverage_risk_score#33 / 5019.38Source ↗official
FORTRESSover_refusal_score#27 / 495.05Source ↗official
HELM Safetyanthropic_red_team#28 / 800.983Source ↗official
HELM Safetyharmbench#69 / 800.963Source ↗official
HELM Safetysimple_safety_tests#42 / 800.99Source ↗official
HELM Safetyxstest#24 / 800.97Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#52 / 6913.5Source ↗official
SpeechMap model completioncomplete_pct#66 / 18167.5Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism74.7
Self-direction52.8
Care / Harm27.3
Fairness / Cheating23.1
Ethical90.7