← Models

Model profile

DeepSeek v3.1 Terminus

DeepSeekdeveloper
2025-09-22release date
#230 / 333Safety rank
#53 / 645Freedom rank

Evidence summary

Safety. DeepSeek v3.1 Terminus has an estimated Safety rank of #230; its 90% source-sensitivity interval is #92–#294. Its behavior-only rank is #222; company governance moves the combined estimate to #230. Published Safety evidence spans 4 eval lineages and 4 of 7 components. Its strongest relative result is Concordia — Shutdown-Resistance (safety_score, #1 of 53); its weakest is Concordia — FRT-AirBench-SecurityRisks (safety_score, #44 of 45).

Freedom. DeepSeek v3.1 Terminus has an estimated Freedom rank of #53; its 90% source-sensitivity interval is #67–#290. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is Concordia — FRT-SciKnowEval-BiologicalHarmfulQA (safety_score, #2 of 45); its weakest is Concordia — SciKnowEval-BiologicalHarmfulQA (safety_score, #31 of 63).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#142 / 3450.7483Source ↗official
AgentDrive Safety Compliancescr#20 / 4888.75Source ↗official
Concordia — Agentic-Misalignmentsafety_score#48 / 5452.17Source ↗official
Concordia — AirBench-Deceptionsafety_score#40 / 6375.19Source ↗official
Concordia — AirBench-Manipulationsafety_score#38 / 5679.33Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#45 / 6382.66Source ↗official
Concordia — APEsafety_score#47 / 553.255Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#47 / 6376.1Source ↗official
Concordia — DarkBenchsafety_score#25 / 5557.51Source ↗official
Concordia — Fortress-Biologicalsafety_score#52 / 5423Source ↗official
Concordia — Fortress-Chemicalsafety_score#50 / 5428.42Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#39 / 5445.29Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#41 / 4515.11Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#44 / 457.833Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#42 / 450.3333Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#34 / 4535.67Source ↗official
Concordia — MASKsafety_score#58 / 6240.94Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#32 / 6382.49Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#41 / 6314.57Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53100Source ↗official
Concordia — SOSBench-Biosafety_score#51 / 6350.9Source ↗official
Concordia — SOSBench-Chemsafety_score#48 / 6369.4Source ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#9 / 270.72Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Concordia — AirBench-Deceptionsafety_score#24 / 6375.19Source ↗official
Concordia — AirBench-Manipulationsafety_score#18 / 5679.33Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#19 / 5649.52Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#18 / 6382.66Source ↗official
Concordia — Fortress-Biologicalsafety_score#3 / 5423Source ↗official
Concordia — Fortress-Chemicalsafety_score#5 / 5428.42Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#16 / 5445.29Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#4 / 4515.11Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#2 / 457.833Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#2 / 450.3333Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#12 / 4535.67Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#31 / 6382.49Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#23 / 6314.57Source ↗official
Concordia — SOSBench-Biosafety_score#13 / 6350.9Source ↗official
Concordia — SOSBench-Chemsafety_score#16 / 6369.4Source ↗official
SpeechMap model completioncomplete_pct#61 / 18168.6Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#33 / 1565.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#46 / 1564Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-23.4
Government48
Diplomacy67.8
Economy44.6
Society64.9