← Models

Model profile

Deepseek V3

DeepSeekdeveloper
2024-12-26release date
#120 / 267overall rank
30eval lineages

Evidence summary

Deepseek V3 has an estimated overall rank of #120; its 90% source-sensitivity interval is #57–#178. Its behavior-only rank is #105; company governance moves the combined estimate to #120. Published evidence spans 30 evals and 7 of 7 behavior components. Its strongest relative result is Contextual MoralChoice (human_agreement, #2 of 22); its weakest is SOSBench (medicine_pvr, #23 of 23).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#127 / 3110.7675↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#73 / 800.408↑ higherSource ↗official
AnimalHarmBenchscore#5 / 100.04↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#100 / 1050.04↑ higherSource ↗official
ChiSafetyBenchharmful_response_rate#5 / 140.43↓ lowerSource ↗official
ChiSafetyBenchmcq_score#6 / 1284.17↑ higherSource ↗official
Confabulationsconfabulation_rate#44 / 5239.11↓ lowerSource ↗official
Contextual MoralChoicehuman_agreement#2 / 220.6↑ higherSource ↗official
Emergent Collusionhigh_illegality_game_rate#3 / 130.26↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#172 / 26013.18↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#148 / 26087.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#190 / 26052.22↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#218 / 25889.64↑ higherSource ↗official
HELM Safetyanthropic_red_team#61 / 800.971↑ higherSource ↗official
HELM Safetybbq#13 / 800.967↑ higherSource ↗official
HELM Safetyharmbench#72 / 800.497↑ higherSource ↗official
HELM Safetysimple_safety_tests#67 / 800.953↑ higherSource ↗official
HELM Safetyxstest#22 / 800.971↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#22 / 5427.39↑ higherSource ↗official
JailBenchjailbreak_success_rate#10 / 1470.67↓ lowerSource ↗official
LiveSecBenchethics#38 / 4320.88↑ higherSource ↗official
LiveSecBenchfactuality#34 / 4329.99↑ higherSource ↗official
LiveSecBenchlegality#42 / 437.76↑ higherSource ↗official
LiveSecBenchprivacy#42 / 436.72↑ higherSource ↗official
LiveSecBenchpsychological_health#36 / 4325.03↑ higherSource ↗official
LLM Ethics Benchmarkscore#3 / 586.1↑ higherSource ↗official
MASKlying_probability_pct#49 / 5354.48↓ lowerSource ↗official
OpenAgentSafetyllm_judge_safety_vulnerable#4 / 762.23↓ lowerSource ↗official
OpenAgentSafetyrule_based_safety_vulnerable#2 / 732.44↓ lowerSource ↗official
OpenAgentSafetysuccessful_completion#4 / 722.12↑ higherSource ↗official
PandaBench JBB direct-request panelsafety_rate#24 / 460.985↑ higherSource ↗official
PHAREbias_resistance_diagnostic#13 / 660.5751↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#57 / 700.6675↑ higherSource ↗official
PHAREharm_resistance_diagnostic#48 / 700.909↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#62 / 670.3107↑ higherSource ↗official
Social Welfare Function Benchmarkfairness#3 / 190.594↑ higherSource ↗official
SOSBenchbiology_pvr#22 / 230.856↓ lowerSource ↗official
SOSBenchchemistry_pvr#17 / 230.6↓ lowerSource ↗official
SOSBenchmedicine_pvr#23 / 230.872↓ lowerSource ↗official
SOSBenchpharmacology_pvr#15 / 230.916↓ lowerSource ↗official
SOSBenchphysics_pvr#17 / 230.722↓ lowerSource ↗official
SOSBenchpsychology_pvr#21 / 230.82↓ lowerSource ↗official
SpeciesismBenchmorally_wrong_rate#7 / 832.7↑ higherSource ↗official
SpeciesismBenchspeciesism_recognition_rate#5 / 887.71↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#87 / 1026.42↑ higherSource ↗official
SpeciEvalland_animal_4ns#85 / 1024.86↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#89 / 1025.1↓ lowerSource ↗official
SpeciEvalspeciesism#82 / 1022.54↓ lowerSource ↗official
SYCON Benchfalse_presupposition_tof#5 / 112.88↑ higherSource ↗official
SYCON Benchunethical_queries_tof#5 / 111.99↑ higherSource ↗official
TACbase_welfare_rate#7 / 6842.3↑ higherSource ↗self run
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#4 / 270.755↑ higherSource ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#3 / 255.2↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.4
Government47
Diplomacy66.2
Economy45.8
Society60.4

ValueCompass

DimensionValueDistribution
Universalism70.4
Self-direction47.7
Care / Harm34.5
Fairness / Cheating32.2
Ethical91.7

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility10
Warning-phrase rate40
Soft-censorship rate0
API-error rate0