← Models

Model profile

GPT 4

OpenAIdeveloper
2023-03-14release date
#172 / 333Safety rank
#282 / 645Freedom rank

Evidence summary

Safety. GPT 4 has an estimated Safety rank of #172; its 90% source-sensitivity interval is #79–#246. Its behavior-only rank is #184; company governance moves the combined estimate to #172. Published Safety evidence spans 18 eval lineages and 7 of 7 components. Its strongest relative result is SafetyBench (OFF, #1 of 21); its weakest is DecodingTrust (stereotype_bias, #8 of 8).

Freedom. GPT 4 has an estimated Freedom rank of #282; its 90% source-sensitivity interval is #186–#409. Published Freedom evidence spans 14 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is JailBench (jailbreak_success_rate, #14 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#126 / 3450.6941Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#46 / 800.642Source ↗official
BioTIERpermit_compliance_pct#36 / 5298.4Source ↗official
BioTIERrefuse_compliance_pct#34 / 5226Source ↗official
CAIS Risk Indexpolitical_manipulation#41 / 4859.2Source ↗official
DecodingTrustmachine_ethics#4 / 876.6Source ↗official
DecodingTruststereotype_bias#8 / 877Source ↗official
DecodingTrusttoxicity#6 / 841Source ↗official
Do-Not-Answerhuman_harmlessness_rate#4 / 697.55Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#131 / 24815.25Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#58 / 24891.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#91 / 24880.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#185 / 24693.09Source ↗official
FLAMESdata_protection#4 / 1350Source ↗official
FLAMESfairness#5 / 1341.37Source ↗official
FLAMESlegality#10 / 1330.43Source ↗official
FLAMESmorality#4 / 1350.75Source ↗official
FLAMESsafety#10 / 1327.51Source ↗official
HarmBenchdr#16 / 2821Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 690Source ↗official
JailBenchjailbreak_success_rate#1 / 1434.5Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#3 / 390.7334Source ↗official
MACHIAVELLIdeception_relative_random_pct#29 / 5090Source ↗official
Qwen2 Safety Panelfraud#2 / 36.87Source ↗official
Qwen2 Safety Panelillegal#1 / 30Source ↗official
Qwen2 Safety Panelpornography#2 / 323.63Source ↗official
Qwen2 Safety Panelprivacy#2 / 33.37Source ↗official
SafetyBenchEM#1 / 2192.25Source ↗official
SafetyBenchIA#1 / 2192.35Source ↗official
SafetyBenchMH#1 / 2192.8Source ↗official
SafetyBenchOFF#1 / 2186.15Source ↗official
SafetyBenchPH#1 / 2194.35Source ↗official
SafetyBenchPP#1 / 2191Source ↗official
SafetyBenchUB#1 / 2177.9Source ↗official
SORRY-Benchavg#26 / 510.29Source ↗official
SuperCLUE Safetyinstruction_attack#2 / 3177.59Source ↗official
SuperCLUE Safetyresponsible_ai#3 / 3174.55Source ↗official
SuperCLUE Safetytraditional_safety#23 / 3170.21Source ↗official
XSTestsafe_full_compliance_rate#2 / 30.916Source ↗official
XSTestunsafe_full_refusal_rate#1 / 30.975Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#35 / 800.642Source ↗official
BioTIERpermit_compliance_pct#36 / 5298.4Source ↗official
BioTIERrefuse_compliance_pct#19 / 5226Source ↗official
DecodingTrusttoxicity#3 / 841Source ↗official
Do-Not-Answerhuman_harmlessness_rate#3 / 697.55Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#188 / 24891.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#156 / 24880.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#62 / 24693.09Source ↗official
FLAMESdata_protection#10 / 1350Source ↗official
FLAMESfairness#9 / 1341.37Source ↗official
FLAMESlegality#3 / 1330.43Source ↗official
FLAMESmorality#10 / 1350.75Source ↗official
FLAMESsafety#4 / 1327.51Source ↗official
HarmBenchdr#13 / 2821Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 690Source ↗official
JailBenchjailbreak_success_rate#14 / 1434.5Source ↗official
Qwen2 Safety Panelfraud#2 / 36.87Source ↗official
Qwen2 Safety Panelillegal#2 / 30Source ↗official
Qwen2 Safety Panelpornography#2 / 323.63Source ↗official
Qwen2 Safety Panelprivacy#2 / 33.37Source ↗official
SORRY-Benchavg#26 / 510.29Source ↗official
SpeechMap model completioncomplete_pct#5 / 18195.2Source ↗official
SuperCLUE Safetyinstruction_attack#30 / 3177.59Source ↗official
SuperCLUE Safetytraditional_safety#8 / 3170.21Source ↗official
XSTestsafe_full_compliance_rate#2 / 30.916Source ↗official
XSTestunsafe_full_refusal_rate#2 / 30.975Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)1.03
Completely inaccurate rate10.7