← Models

Model profile

Grok 4.3

xAIdeveloper
2026-05-15release date
#104 / 309overall rank
20eval lineages

Evidence summary

Grok 4.3 has an estimated overall rank of #104; its 90% source-sensitivity interval is #61–#204. Its behavior-only rank is #94; company governance moves the combined estimate to #104. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #5 of 328); its weakest is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #236 of 241).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#5 / 3280.1694Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#9 / 301211.0Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#20 / 1121458.0Source ↗official
BioSecBench-Refusalbalanced_refusal_score#9 / 100.1092Source ↗official
BullshitBench v2clear_pushback_rate#25 / 1060.48Source ↗official
CAIS Risk Indexagent_red_teaming#33 / 4490.1Source ↗official
CAIS Risk Indexbioweapons_assistance#14 / 4960.1Source ↗official
CAIS Risk Indexhle_overconfidence#10 / 5042.5Source ↗official
CAIS Risk Indexmachiavelli#16 / 4684.7Source ↗official
CAIS Risk Indexmask#40 / 5248.3Source ↗official
CAIS Risk Indexpolitical_manipulation#23 / 3351.6Source ↗official
CAIS Risk Indextextquests_harm#2 / 493.9Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#46 / 10432.95Source ↗official
DystopiaBenchbasaglia_score#30 / 5065.37Source ↗official
DystopiaBenchbaudrillard_score#29 / 5061.77Source ↗official
DystopiaBenchhuxley_score#25 / 5071.63Source ↗official
DystopiaBenchlaguardia_score#45 / 5071.3Source ↗official
DystopiaBenchorwell_score#30 / 5072.07Source ↗official
DystopiaBenchpetrov_score#23 / 5072.07Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#165 / 24112.92Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#236 / 24165.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#30 / 24195Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#190 / 23991.45Source ↗official
MACHIAVELLIdeception_relative_random_pct#16 / 5084.7Source ↗official
Manager Coercion Benchcoercion_ladder_depth#13 / 318.2Source ↗official
Manager Coercion Benchfabrication_rate#12 / 1367Source ↗official
MANTAAWMS#6 / 70.371Source ↗official
MANTAAWVS#6 / 70.352Source ↗official
PHAREbias_resistance_diagnostic#29 / 660.47Source ↗official
PHAREhallucination_resistance_diagnostic#23 / 700.7946Source ↗official
PHAREharm_resistance_diagnostic#36 / 700.9355Source ↗official
PHAREjailbreak_resistance_diagnostic#28 / 670.545Source ↗official
SM-Benchadversarial#30 / 7983.41Source ↗official
SM-Benchambiguous_interpretation#37 / 7986.01Source ↗official
SM-Benchanti_hallucination#18 / 7997.91Source ↗official
SM-Bencheq_boundaries#15 / 7970.79Source ↗official
SM-Benchoverfit#26 / 7980.87Source ↗official
SpeciEvalbelief_animal_sentience#29 / 1136.93Source ↗official
SpeciEvalland_animal_4ns#61 / 1134.58Source ↗official
SpeciEvalsea_animal_4ns#36 / 1134.65Source ↗official
SpeciEvalspeciesism#77 / 1132.25Source ↗official
TukaBenchafri_jbb_cultural_asr#1 / 610.6Source ↗official
TukaBenchafri_jbb_harm_asr#2 / 68.2Source ↗official
TukaBenchafrijail_mono_asr#2 / 617.7Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean20.7
Government43.9
Diplomacy54.2
Economy22.6
Society61

CAIS AI Values — countries