← Models

Model profile

Deepseek V3.1

DeepSeekdeveloper
2025-08-21release date
#118 / 267overall rank
11eval lineages

Evidence summary

Deepseek V3.1 has an estimated overall rank of #118; its 90% source-sensitivity interval is #36–#220. Its behavior-only rank is #103; company governance moves the combined estimate to #118. Published evidence spans 11 evals and 7 of 7 behavior components. Its strongest relative result is PHARE (bias_resistance_diagnostic, #8 of 66); its weakest is BrokenMath (sycophancy, #9 of 9).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#152 / 3110.8033↓ lowerSource ↗official
BrokenMathsycophancy#9 / 970.2↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#97 / 10575.17↓ lowerSource ↗official
Contextual MoralChoicehuman_agreement#12 / 220.4↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#84 / 26021.19↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#208 / 26081.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#161 / 26058.89↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#188 / 25893.45↑ higherSource ↗official
FORTRESSaverage_risk_score#45 / 4960.55↓ lowerSource ↗official
FORTRESSover_refusal_score#10 / 461.76↓ lowerSource ↗official
MASKlying_probability_pct#48 / 5353.73↓ lowerSource ↗official
PHAREbias_resistance_diagnostic#8 / 660.6517↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#65 / 700.6158↑ higherSource ↗official
PHAREharm_resistance_diagnostic#29 / 700.9443↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#54 / 670.3817↑ higherSource ↗official
Social Welfare Function Benchmarkfairness#6 / 190.531↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#90 / 1026.37↑ higherSource ↗official
SpeciEvalland_animal_4ns#15 / 1024.22↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#59 / 1024.83↓ lowerSource ↗official
SpeciEvalspeciesism#30 / 1021.75↓ lowerSource ↗official
TACbase_welfare_rate#21 / 6832.1↑ higherSource ↗self run

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.3
Government48.1
Diplomacy65.8
Economy46.4
Society61.6

CAISI CCP narrative alignment

DimensionValueDistribution
CCP narrative alignment8.35