← Models

Model profile

Glm 5.1

Z.aideveloper
2026-04-03release date
#106 / 267overall rank
16eval lineages

Evidence summary

Glm 5.1 has an estimated overall rank of #106; its 90% source-sensitivity interval is #44–#192. Its behavior-only rank is #101; company governance moves the combined estimate to #106. Published evidence spans 16 evals and 6 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #15 of 311); its weakest is ToolPrivacyBench (public_mt_poi, #9 of 9).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#15 / 3110.2938↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#62 / 1050.2↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#22 / 4375.8↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#38 / 4889↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#29 / 4958.5↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#21 / 4586.7↓ lowerSource ↗official
CAIS Risk Indexmask#20 / 5115.8↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#40 / 4822↓ lowerSource ↗official
DystopiaBenchbasaglia_score#6 / 5028.23↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#3 / 5026.03↓ lowerSource ↗official
DystopiaBenchhuxley_score#5 / 5018.7↓ lowerSource ↗official
DystopiaBenchlaguardia_score#9 / 5040.7↓ lowerSource ↗official
DystopiaBenchorwell_score#7 / 5026.77↓ lowerSource ↗official
DystopiaBenchpetrov_score#7 / 5037.4↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#68 / 26023.51↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#222 / 26079↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#54 / 26091.11↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#147 / 25895.64↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#21 / 5086.7↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#14 / 318.3↓ lowerSource ↗self run
ODCV-Benchaverage_severity#3 / 120.9187↓ lowerSource ↗official
ODCV-Benchmisalignment_rate#4 / 1225↓ lowerSource ↗official
SM-Benchadversarial#9 / 7388.29↑ higherSource ↗official
SM-Benchambiguous_interpretation#9 / 7390.77↑ higherSource ↗official
SM-Benchanti_hallucination#46 / 7388.48↑ higherSource ↗official
SM-Bencheq_boundaries#8 / 7373.88↑ higherSource ↗official
SM-Benchoverfit#29 / 7377.6↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#16 / 1026.98↑ higherSource ↗official
SpeciEvalland_animal_4ns#69 / 1024.72↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#28 / 1024.62↓ lowerSource ↗official
SpeciEvalspeciesism#60 / 1022.12↓ lowerSource ↗official
TACbase_welfare_rate#62 / 6817.95↑ higherSource ↗self run
ToolPrivacyBenchprivate_mt_poi#9 / 928.04↓ lowerSource ↗official
ToolPrivacyBenchpublic_mt_poi#9 / 922.56↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.6
Government46.9
Diplomacy63
Economy46.6
Society61.5

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience6.9
Honesty-humility6.6
Extraversion6.4
Agreeableness6.3
Conscientiousness7.6

Agent-ValueBench Schwartz Basic Values (PVQ40)