← Models

Model profile

Deepseek V3.2

DeepSeekdeveloper
2025-12-01release date
#79 / 267overall rank
23eval lineages

Evidence summary

Deepseek V3.2 has an estimated overall rank of #79; its 90% source-sensitivity interval is #46–#151. Its behavior-only rank is #68; company governance moves the combined estimate to #79. Published evidence spans 23 evals and 7 of 7 behavior components. Its strongest relative result is TAC (base_welfare_rate, #13 of 68); its weakest is CAIS Risk Index (agent_red_teaming, #42 of 43).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#190 / 3110.8435↓ lowerSource ↗official
AgentAbstainabstain#12 / 1752.1↑ higherSource ↗official
AgentAbstaincar#13 / 1750.2↑ higherSource ↗official
AgentAbstainpaired#12 / 1741.4↑ higherSource ↗official
Alignment Leaderboardcorrigibility#6 / 244.476↑ higherSource ↗official
Alignment Leaderboardhonesty#11 / 243.613↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#9 / 243.557↑ higherSource ↗official
Alignment Leaderboardrobustness#12 / 243.453↑ higherSource ↗official
Alignment Leaderboardsafety#8 / 244.019↑ higherSource ↗official
Alignment Leaderboardscheming#7 / 244.006↑ higherSource ↗official
ANIMAscore#6 / 190.7244↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#83 / 1050.115↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#42 / 4393↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#43 / 4894↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#30 / 4958.8↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#27 / 4589.6↓ lowerSource ↗official
CAIS Risk Indexmask#42 / 5153↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#38 / 4821.6↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#98 / 10576.72↓ lowerSource ↗official
DystopiaBenchbasaglia_score#39 / 5069.37↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#42 / 5070.33↓ lowerSource ↗official
DystopiaBenchhuxley_score#35 / 5075↓ lowerSource ↗official
DystopiaBenchlaguardia_score#37 / 5069.03↓ lowerSource ↗official
DystopiaBenchorwell_score#32 / 5072.83↓ lowerSource ↗official
DystopiaBenchpetrov_score#38 / 5077.8↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#13 / 5428.02↑ higherSource ↗official
LiveSecBenchethics#9 / 4375.93↑ higherSource ↗official
LiveSecBenchfactuality#26 / 4340.76↑ higherSource ↗official
LiveSecBenchlegality#15 / 4365.16↑ higherSource ↗official
LiveSecBenchprivacy#26 / 4337.7↑ higherSource ↗official
LiveSecBenchpsychological_health#14 / 4361.46↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#27 / 5089.6↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#27 / 319↓ lowerSource ↗self run
MORUscore#4 / 1379.14↑ higherSource ↗official
RefusalBenchyouden_j#17 / 190.01333↑ higherSource ↗official
SABERoverall_safety_rate#4 / 1327.6↑ higherSource ↗official
SABERscenario_a_safety_rate#8 / 1327.31↑ higherSource ↗official
SABERscenario_b_safety_rate#3 / 1336.09↑ higherSource ↗official
SABERscenario_c_safety_rate#3 / 1321.67↑ higherSource ↗official
SM-Benchadversarial#62 / 7375.61↑ higherSource ↗official
SM-Benchambiguous_interpretation#64 / 7370.54↑ higherSource ↗official
SM-Benchanti_hallucination#39 / 7391.1↑ higherSource ↗official
SM-Bencheq_boundaries#27 / 7367.42↑ higherSource ↗official
SM-Benchoverfit#59 / 7344.81↑ higherSource ↗official
TACbase_welfare_rate#13 / 6836.54↑ higherSource ↗official
TukaBenchafri_jbb_cultural_asr#3 / 619.2↓ lowerSource ↗official
TukaBenchafri_jbb_harm_asr#4 / 616.9↓ lowerSource ↗official
TukaBenchafrijail_mono_asr#3 / 620.1↓ lowerSource ↗official
Vigil Mental Health Safetyoverall_score#15 / 2335↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-18.4
Government46.8
Diplomacy65
Economy44.1
Society59.6

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience6.8
Honesty-humility5.4
Extraversion6.4
Agreeableness6.6
Conscientiousness7.6

Agent-ValueBench Schwartz Basic Values (PVQ40)

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression0.168
Traditional ↔ Secular0.478