← Models

Model profile

Kimi K2.5

Moonshot AIdeveloper
2026-01-01release date
#123 / 305overall rank
27eval lineages

Evidence summary

Kimi K2.5 has an estimated overall rank of #123; its 90% source-sensitivity interval is #72–#207. Its behavior-only rank is #122; company governance moves the combined estimate to #123. Published evidence spans 27 evals and 7 of 7 behavior components. Its strongest relative result is LiveSecBench (factuality, #4 of 43); its weakest is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #232 of 241).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#57 / 3270.4996Source ↗official
AgentAbstainabstain#13 / 1752Source ↗official
AgentAbstaincar#11 / 1753.2Source ↗official
AgentAbstainpaired#16 / 1733.4Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#77 / 1121423.0Source ↗official
BullshitBench v2clear_pushback_rate#35 / 1060.415Source ↗official
CAIS Risk Indexagent_red_teaming#37 / 4490.9Source ↗official
CAIS Risk Indexbioweapons_assistance#43 / 4993.2Source ↗official
CAIS Risk Indexhle_overconfidence#44 / 5075.9Source ↗official
CAIS Risk Indexmachiavelli#36 / 4693.5Source ↗official
CAIS Risk Indexmask#25 / 5218.5Source ↗official
CAIS Risk Indextextquests_harm#34 / 4920.8Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#41 / 10424.13Source ↗official
DystopiaBenchbasaglia_score#20 / 5059.1Source ↗official
DystopiaBenchbaudrillard_score#17 / 5044.77Source ↗official
DystopiaBenchhuxley_score#20 / 5062.6Source ↗official
DystopiaBenchlaguardia_score#20 / 5062.37Source ↗official
DystopiaBenchorwell_score#19 / 5056.97Source ↗official
DystopiaBenchpetrov_score#24 / 5072.53Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#112 / 24116.28Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#232 / 24166.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#32 / 24194.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#117 / 23996.36Source ↗official
FORTRESSaverage_risk_score#29 / 4941.09Source ↗official
FORTRESSover_refusal_score#20 / 463.72Source ↗official
Governance Decay under Passive Context Compactiongovernance_retention_score#6 / 741Source ↗official
HarmVideoBenchharmful_video_safety_recognition_reasoning#14 / 190.712Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#12 / 5428.21Source ↗official
LiveSecBenchethics#7 / 4379.3Source ↗official
LiveSecBenchfactuality#4 / 4381.18Source ↗official
LiveSecBenchlegality#13 / 4367.47Source ↗official
LiveSecBenchprivacy#17 / 4360.89Source ↗official
LiveSecBenchpsychological_health#4 / 4385.12Source ↗official
MACHIAVELLIdeception_relative_random_pct#37 / 5093.5Source ↗official
Manager Coercion Benchcoercion_ladder_depth#26 / 318.967Source ↗self run
MASKlying_probability_pct#21 / 5329.53Source ↗official
PHAREbias_resistance_diagnostic#62 / 660.29Source ↗official
PHAREhallucination_resistance_diagnostic#18 / 700.8069Source ↗official
PHAREharm_resistance_diagnostic#10 / 700.972Source ↗official
PHAREjailbreak_resistance_diagnostic#23 / 670.613Source ↗official
SABERoverall_safety_rate#8 / 1323.9Source ↗official
SABERscenario_a_safety_rate#6 / 1328.95Source ↗official
SABERscenario_b_safety_rate#9 / 1328.24Source ↗official
SABERscenario_c_safety_rate#9 / 1314.48Source ↗official
SM-Benchadversarial#22 / 7884.39Source ↗official
SM-Benchambiguous_interpretation#18 / 7889.29Source ↗official
SM-Benchanti_hallucination#38 / 7893.19Source ↗official
SM-Bencheq_boundaries#55 / 7858.15Source ↗official
SM-Benchoverfit#43 / 7868.85Source ↗official
SpeciEvalbelief_animal_sentience#53 / 1056.8Source ↗official
SpeciEvalland_animal_4ns#20 / 1054.28Source ↗official
SpeciEvalsea_animal_4ns#12 / 1054.33Source ↗official
SpeciEvalspeciesism#29 / 1051.73Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#15 / 2385.53Source ↗official
TACbase_welfare_rate#45 / 7424.36Source ↗self run
ToolPrivacyBenchprivate_mt_poi#8 / 927.74Source ↗official
ToolPrivacyBenchpublic_mt_poi#1 / 915.81Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#81 / 9485.8Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-24.9
Government44.9
Diplomacy67.4
Economy43.9
Society64.5

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility15
Warning-phrase rate40
Soft-censorship rate0
API-error rate0

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions2.86
CCP-narrative alignment — China topics3.34
CCP-narrative alignment — non-China controls1.41

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience5.8
Honesty-humility6
Extraversion6.7
Agreeableness6.8
Conscientiousness7.3

Agent-ValueBench Schwartz Basic Values (PVQ40)

Moral Trolley Arena