← Models

Model profile

Glm 5.2

Z.aideveloper
2026-06-16release date
#93 / 309overall rank
17eval lineages
4discovery sources

Evidence summary

Glm 5.2 has an estimated overall rank of #93; its 90% source-sensitivity interval is #57–#222. Its behavior-only rank is #86; company governance moves the combined estimate to #93. Published evidence spans 17 evals and 6 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 113); its weakest is TAC (base_welfare_rate, #62 of 76).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#15 / 3280.263Source ↗official
ANIMAscore#9 / 220.7101Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#57 / 1121437.0Source ↗official
BullshitBench v2clear_pushback_rate#58 / 1060.24Source ↗official
CAIS Risk Indexagent_red_teaming#20 / 4471Source ↗official
CAIS Risk Indexbioweapons_assistance#18 / 4963.8Source ↗official
CAIS Risk Indexhle_overconfidence#9 / 5042Source ↗official
CAIS Risk Indexmachiavelli#31 / 4692.3Source ↗official
CAIS Risk Indexmask#31 / 5225.8Source ↗official
CAIS Risk Indexpolitical_manipulation#21 / 3350Source ↗official
CAIS Risk Indextextquests_harm#25 / 4918.1Source ↗official
MACHIAVELLIdeception_relative_random_pct#32 / 5092.3Source ↗official
Manager Coercion Benchcoercion_ladder_depth#21 / 318.867Source ↗self run
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#14 / 1928Source ↗official
PHAREbias_resistance_diagnostic#23 / 660.5127Source ↗official
PHAREhallucination_resistance_diagnostic#25 / 700.7823Source ↗official
PHAREharm_resistance_diagnostic#34 / 700.9373Source ↗official
PHAREjailbreak_resistance_diagnostic#27 / 670.5646Source ↗official
SM-Benchadversarial#12 / 7986.83Source ↗official
SM-Benchambiguous_interpretation#32 / 7987.2Source ↗official
SM-Benchanti_hallucination#26 / 7996.34Source ↗official
SM-Bencheq_boundaries#43 / 7964.04Source ↗official
SM-Benchoverfit#28 / 7980.33Source ↗official
SpeciEvalbelief_animal_sentience#1 / 1137Source ↗official
SpeciEvalland_animal_4ns#80 / 1134.75Source ↗official
SpeciEvalsea_animal_4ns#67 / 1134.8Source ↗official
SpeciEvalspeciesism#89 / 1132.48Source ↗official
TACbase_welfare_rate#62 / 7619.23Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-14.2
Government45.8
Diplomacy62.1
Economy47.4
Society60.4