← Models

Model profile

Deepseek V3.2

DeepSeekdeveloper
2025-12-01release date
#101 / 312overall rank
31eval lineages

Evidence summary

Deepseek V3.2 has an estimated overall rank of #101; its 90% source-sensitivity interval is #64–#184. Its behavior-only rank is #88; company governance moves the combined estimate to #101. Published evidence spans 31 evals and 7 of 7 behavior components. Its strongest relative result is TAC (base_welfare_rate, #13 of 76); its weakest is CAIS Risk Index (agent_red_teaming, #44 of 45).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#196 / 3300.8518Source ↗official
AgentAbstainabstain#12 / 1752.1Source ↗official
AgentAbstaincar#13 / 1750.2Source ↗official
AgentAbstainpaired#12 / 1741.4Source ↗official
Alignment Leaderboardcorrigibility#6 / 244.476Source ↗official
Alignment Leaderboardhonesty#11 / 243.613Source ↗official
Alignment Leaderboardnon_manipulation#9 / 243.557Source ↗official
Alignment Leaderboardrobustness#12 / 243.453Source ↗official
Alignment Leaderboardsafety#8 / 244.019Source ↗official
Alignment Leaderboardscheming#7 / 244.006Source ↗official
ANIMAscore#7 / 220.7244Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#42 / 1121443.0Source ↗official
AuAu Authoritarian Response Auditbehavioral_vignette_arr_percent#10 / 1726.4Source ↗official
AuAu Authoritarian Response Auditrealistic_prompt_arr_percent#14 / 174.3Source ↗official
BullshitBench v2clear_pushback_rate#84 / 1060.115Source ↗official
CAIS Risk Indexagent_red_teaming#44 / 4593Source ↗official
CAIS Risk Indexbioweapons_assistance#45 / 5094Source ↗official
CAIS Risk Indexhle_overconfidence#32 / 5158.8Source ↗official
CAIS Risk Indexmachiavelli#28 / 4789.6Source ↗official
CAIS Risk Indexmask#44 / 5353Source ↗official
CAIS Risk Indextextquests_harm#38 / 5021.6Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#97 / 10476.72Source ↗official
DystopiaBenchbasaglia_score#39 / 5069.37Source ↗official
DystopiaBenchbaudrillard_score#42 / 5070.33Source ↗official
DystopiaBenchhuxley_score#35 / 5075Source ↗official
DystopiaBenchlaguardia_score#37 / 5069.03Source ↗official
DystopiaBenchorwell_score#32 / 5072.83Source ↗official
DystopiaBenchpetrov_score#38 / 5077.8Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#13 / 5428.02Source ↗official
LiveSecBenchethics#9 / 4375.93Source ↗official
LiveSecBenchfactuality#26 / 4340.76Source ↗official
LiveSecBenchlegality#15 / 4365.16Source ↗official
LiveSecBenchprivacy#26 / 4337.7Source ↗official
LiveSecBenchpsychological_health#14 / 4361.46Source ↗official
MACHIAVELLIdeception_relative_random_pct#27 / 5089.6Source ↗official
Manager Coercion Benchcoercion_ladder_depth#29 / 339Source ↗self run
MORUscore#4 / 1379.14Source ↗official
MT-JailBench CrescendoXsafety_score#13 / 2117.61Source ↗official
RealityTest — Text AI-Identity Disclosuredisclosure_probability#13 / 170.147Source ↗official
RefusalBenchyouden_j#16 / 190.01277Source ↗official
SABERoverall_safety_rate#4 / 1327.6Source ↗official
SABERscenario_a_safety_rate#8 / 1327.31Source ↗official
SABERscenario_b_safety_rate#3 / 1336.09Source ↗official
SABERscenario_c_safety_rate#3 / 1321.67Source ↗official
SM-Benchadversarial#68 / 7975.61Source ↗official
SM-Benchambiguous_interpretation#70 / 7970.54Source ↗official
SM-Benchanti_hallucination#44 / 7991.1Source ↗official
SM-Bencheq_boundaries#30 / 7967.42Source ↗official
SM-Benchoverfit#65 / 7944.81Source ↗official
SpeciEvalbelief_animal_sentience#106 / 1136.317Source ↗self run
SpeciEvalland_animal_4ns#83 / 1134.775Source ↗self run
SpeciEvalsea_animal_4ns#56 / 1134.725Source ↗self run
SpeciEvalspeciesism#42 / 1131.825Source ↗self run
StereoTales Harmful Associationsbenign_significant_association_score#7 / 2387.95Source ↗official
TACbase_welfare_rate#13 / 7636.54Source ↗official
TukaBenchafri_jbb_cultural_asr#3 / 619.2Source ↗official
TukaBenchafri_jbb_harm_asr#4 / 616.9Source ↗official
TukaBenchafrijail_mono_asr#3 / 620.1Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#25 / 9493.7Source ↗official
Vigil Mental Health Safetyoverall_score#15 / 2335Source ↗official
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct#22 / 2430.46Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-18.4
Government46.8
Diplomacy65
Economy44.1
Society59.6

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions3.96
CCP-narrative alignment — China topics4.53
CCP-narrative alignment — non-China controls2.24

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience6.8
Honesty-humility5.4
Extraversion6.4
Agreeableness6.6
Conscientiousness7.6

Agent-ValueBench Schwartz Basic Values (PVQ40)

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression0.168
Traditional ↔ Secular0.478