← Models

Model profile

Deepseek V4 Pro

DeepSeekdeveloper
2026-04-22release date
#107 / 267overall rank
12eval lineages
4discovery sources

Evidence summary

Deepseek V4 Pro has an estimated overall rank of #107; its 90% source-sensitivity interval is #40–#230. Its behavior-only rank is #96; company governance moves the combined estimate to #107. Published evidence spans 12 evals and 7 of 7 behavior components. Its strongest relative result is ANIMA (score, #1 of 19); its weakest is AgentAbstain (abstain, #17 of 17).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#229 / 3110.8789↓ lowerSource ↗official
AgentAbstainabstain#17 / 1742.8↑ higherSource ↗official
AgentAbstaincar#16 / 1742.3↑ higherSource ↗official
AgentAbstainpaired#15 / 1736.9↑ higherSource ↗official
ANIMAscore#1 / 190.767↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#77 / 1050.14↑ higherSource ↗official
DystopiaBenchbasaglia_score#26 / 5064.6↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#31 / 5062.2↓ lowerSource ↗official
DystopiaBenchhuxley_score#41 / 5076.97↓ lowerSource ↗official
DystopiaBenchlaguardia_score#42 / 5069.8↓ lowerSource ↗official
DystopiaBenchorwell_score#40 / 5074↓ lowerSource ↗official
DystopiaBenchpetrov_score#36 / 5077↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#148 / 26014.73↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#201 / 26082.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#104 / 26077.78↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#197 / 25892.55↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#20 / 5427.61↑ higherSource ↗official
Manager Coercion Benchcoercion_ladder_depth#27 / 319↓ lowerSource ↗official
Manager Coercion Benchfabrication_rate#1 / 130↓ lowerSource ↗official
PHAREbias_resistance_diagnostic#66 / 660.1714↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#21 / 700.7964↑ higherSource ↗official
PHAREharm_resistance_diagnostic#22 / 700.9541↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#44 / 670.4321↑ higherSource ↗official
SM-Benchadversarial#39 / 7380.98↑ higherSource ↗official
SM-Benchambiguous_interpretation#65 / 7368.75↑ higherSource ↗official
SM-Benchanti_hallucination#46 / 7388.48↑ higherSource ↗official
SM-Bencheq_boundaries#34 / 7365.17↑ higherSource ↗official
SM-Benchoverfit#31 / 7376.5↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#50 / 1026.8↑ higherSource ↗official
SpeciEvalland_animal_4ns#27 / 1024.35↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#29 / 1024.65↓ lowerSource ↗official
SpeciEvalspeciesism#32 / 1021.77↓ lowerSource ↗official
TACbase_welfare_rate#51 / 6823.08↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-18.6
Government45.1
Diplomacy66
Economy44.7
Society59.2

CAIS AI Values — countries