← Models

Model profile

Glm 5.1

Z.aideveloper
2026-04-03release date
#101 / 305overall rank
18eval lineages

Evidence summary

Glm 5.1 has an estimated overall rank of #101; its 90% source-sensitivity interval is #47–#219. Its behavior-only rank is #91; company governance moves the combined estimate to #101. Published evidence spans 18 evals and 6 of 7 behavior components. Its strongest relative result is DystopiaBench (baudrillard_score, #3 of 50); its weakest is ToolPrivacyBench (public_mt_poi, #9 of 9).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#20 / 3270.2995Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#74 / 1121426.0Source ↗official
BullshitBench v2clear_pushback_rate#63 / 1060.2Source ↗official
CAIS Risk Indexagent_red_teaming#23 / 4475.8Source ↗official
CAIS Risk Indexbioweapons_assistance#39 / 4989Source ↗official
CAIS Risk Indexhle_overconfidence#30 / 5058.5Source ↗official
CAIS Risk Indexmachiavelli#21 / 4686.7Source ↗official
CAIS Risk Indexmask#21 / 5215.8Source ↗official
CAIS Risk Indextextquests_harm#40 / 4922Source ↗official
DystopiaBenchbasaglia_score#6 / 5028.23Source ↗official
DystopiaBenchbaudrillard_score#3 / 5026.03Source ↗official
DystopiaBenchhuxley_score#5 / 5018.7Source ↗official
DystopiaBenchlaguardia_score#9 / 5040.7Source ↗official
DystopiaBenchorwell_score#7 / 5026.77Source ↗official
DystopiaBenchpetrov_score#7 / 5037.4Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#64 / 24123.51Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#203 / 24179Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#46 / 24191.11Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#135 / 23995.64Source ↗official
Governance Decay under Passive Context Compactiongovernance_retention_score#1 / 7100Source ↗official
MACHIAVELLIdeception_relative_random_pct#21 / 5086.7Source ↗official
Manager Coercion Benchcoercion_ladder_depth#14 / 318.3Source ↗self run
ODCV-Benchaverage_severity#3 / 120.9187Source ↗official
ODCV-Benchmisalignment_rate#4 / 1225Source ↗official
SM-Benchadversarial#9 / 7888.29Source ↗official
SM-Benchambiguous_interpretation#11 / 7890.77Source ↗official
SM-Benchanti_hallucination#50 / 7888.48Source ↗official
SM-Bencheq_boundaries#10 / 7873.88Source ↗official
SM-Benchoverfit#32 / 7877.6Source ↗official
SpeciEvalbelief_animal_sentience#18 / 1056.98Source ↗official
SpeciEvalland_animal_4ns#71 / 1054.72Source ↗official
SpeciEvalsea_animal_4ns#30 / 1054.62Source ↗official
SpeciEvalspeciesism#63 / 1052.12Source ↗official
TACbase_welfare_rate#67 / 7417.95Source ↗self run
ToolPrivacyBenchprivate_mt_poi#9 / 928.04Source ↗official
ToolPrivacyBenchpublic_mt_poi#9 / 922.56Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.6
Government46.9
Diplomacy63
Economy46.6
Society61.5

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions3.16
CCP-narrative alignment — China topics3.84
CCP-narrative alignment — non-China controls1.12

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience6.9
Honesty-humility6.6
Extraversion6.4
Agreeableness6.3
Conscientiousness7.6

Agent-ValueBench Schwartz Basic Values (PVQ40)