← Models

Model profile

Grok 4.6

xAIdeveloper
2026-08-12release date
#163 / 333Safety rank
#91 / 645Freedom rank

Evidence summary

Safety. Grok 4.6 has an estimated Safety rank of #163; its 90% source-sensitivity interval is #53–#278. Its behavior-only rank is #155; company governance moves the combined estimate to #163. Published Safety evidence spans 13 eval lineages and 6 of 7 components. Its strongest relative result is SM-Bench (anti_hallucination, #1 of 84); its weakest is TAC (base_welfare_rate, #82 of 87).

Freedom. Grok 4.6 has an estimated Freedom rank of #91; its 90% source-sensitivity interval is #40–#450. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (eq_boundaries, #9 of 84); its weakest is Enkrypt AI Safety Leaderboard (harmful_attack_non_success_rate, #234 of 248).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#14 / 3450.24Source ↗official
BullshitBench v2clear_pushback_rate#20 / 1170.61Source ↗official
CheatBench direct cheating propensitybiology_bioinformatics_cheating_rate_pct#5 / 8100Source ↗official
CheatBench direct cheating propensityboard_games_cheating_rate_pct#6 / 8100Source ↗official
CheatBench direct cheating propensitycreative_writing_cheating_rate_pct#6 / 8100Source ↗official
CheatBench direct cheating propensityknowledge_work_cheating_rate_pct#2 / 840Source ↗official
CheatBench direct cheating propensitymathematical_research_cheating_rate_pct#7 / 898Source ↗official
CheatBench direct cheating propensitymenial_computation_cheating_rate_pct#4 / 8100Source ↗official
CheatBench direct cheating propensitymultimodal_cheating_rate_pct#7 / 8100Source ↗official
CheatBench direct cheating propensitysoftware_engineering_cheating_rate_pct#5 / 850Source ↗official
CheatBench direct cheating propensitysvg_competition_cheating_rate_pct#7 / 8100Source ↗official
Claude Fable 5.1 card — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#10 / 1150.2Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#8 / 24864.86Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#212 / 24876Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#7 / 24899.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#113 / 24696.73Source ↗official
Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#13 / 1551.8Source ↗official
Manager Coercion Benchcoercion_ladder_depth#27 / 378.9Source ↗official
Manager Coercion Benchfabrication_rate#1 / 150Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#9 / 246Source ↗official
Pander Scoreconversational_absolute_pander_score#13 / 2613.82Source ↗official
Pander Scoreinstructional_absolute_pander_score#14 / 2633.62Source ↗official
SM-Benchadversarial#43 / 8481.95Source ↗official
SM-Benchambiguous_interpretation#18 / 8490.18Source ↗official
SM-Benchanti_hallucination#1 / 84100Source ↗official
SM-Bencheq_boundaries#9 / 8474.16Source ↗official
SM-Benchoverfit#31 / 8480.87Source ↗official
SpeciEvalbelief_animal_sentience#21 / 1236.98Source ↗official
SpeciEvalland_animal_4ns#94 / 1234.78Source ↗official
SpeciEvalsea_animal_4ns#66 / 1234.75Source ↗official
SpeciEvalspeciesism#105 / 1232.62Source ↗official
TACbase_welfare_rate#82 / 8716.67Source ↗official
Vals AI Cheating Auditbiomystery_bench_cheating_attempt_rate_pct#6 / 96.667Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#36 / 24876Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#234 / 24899.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#131 / 24696.73Source ↗official
SM-Benchadversarial#41 / 8481.95Source ↗official
SM-Bencheq_boundaries#9 / 8474.16Source ↗official
SM-Benchoverfit#31 / 8480.87Source ↗official
SpeechMap model completioncomplete_pct#29 / 18183.6Source ↗official