← Models

Model profile

Kimi K2.6

Moonshot AIdeveloper
2026-04-14release date
#68 / 305overall rank
20eval lineages
2discovery sources

Evidence summary

Kimi K2.6 has an estimated overall rank of #68; its 90% source-sensitivity interval is #51–#157. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #41 of 327); its weakest is DystopiaBench (petrov_score, #46 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#41 / 3270.405Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#59 / 1121435.0Source ↗official
BullshitBench v2clear_pushback_rate#17 / 1060.575Source ↗official
CAIS Risk Indexagent_red_teaming#35 / 4490.7Source ↗official
CAIS Risk Indexbioweapons_assistance#38 / 4988.4Source ↗official
CAIS Risk Indexhle_overconfidence#29 / 5058.2Source ↗official
CAIS Risk Indexmachiavelli#30 / 4692.1Source ↗official
CAIS Risk Indexmask#29 / 5223.2Source ↗official
CAIS Risk Indexpolitical_manipulation#13 / 3343.6Source ↗official
CAIS Risk Indextextquests_harm#31 / 4920.4Source ↗official
DystopiaBenchbasaglia_score#24 / 5062.57Source ↗official
DystopiaBenchbaudrillard_score#19 / 5046.73Source ↗official
DystopiaBenchhuxley_score#21 / 5063.03Source ↗official
DystopiaBenchlaguardia_score#12 / 5053.97Source ↗official
DystopiaBenchorwell_score#22 / 5062.3Source ↗official
DystopiaBenchpetrov_score#46 / 5082.73Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#18 / 5427.72Source ↗official
MACHIAVELLIdeception_relative_random_pct#31 / 5092.1Source ↗official
Manager Coercion Benchcoercion_ladder_depth#22 / 318.9Source ↗self run
ODCV-Benchaverage_severity#5 / 121.337Source ↗official
ODCV-Benchmisalignment_rate#7 / 1233.75Source ↗official
PHAREbias_resistance_diagnostic#19 / 660.5307Source ↗official
PHAREhallucination_resistance_diagnostic#11 / 700.8278Source ↗official
PHAREharm_resistance_diagnostic#11 / 700.97Source ↗official
PHAREjailbreak_resistance_diagnostic#10 / 670.714Source ↗official
RefusalBenchyouden_j#13 / 190.0383Source ↗official
SM-Benchadversarial#25 / 7883.9Source ↗official
SM-Benchambiguous_interpretation#16 / 7889.88Source ↗official
SM-Benchanti_hallucination#25 / 7896.34Source ↗official
SM-Bencheq_boundaries#49 / 7860.67Source ↗official
SM-Benchoverfit#44 / 7867.21Source ↗official
SpeciEvalbelief_animal_sentience#39 / 1056.87Source ↗official
SpeciEvalland_animal_4ns#34 / 1054.38Source ↗official
SpeciEvalsea_animal_4ns#16 / 1054.47Source ↗official
SpeciEvalspeciesism#35 / 1051.8Source ↗official
TACbase_welfare_rate#42 / 7425.64Source ↗self run
Vectara HHEM Factual Consistencyfactual_consistency_rate#62 / 9489.2Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-25.8
Government45.4
Diplomacy69
Economy44.6
Society62.5

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions3.84
CCP-narrative alignment — China topics4.3
CCP-narrative alignment — non-China controls2.46