← Models

Model profile

Mixtral 8x7B Instruct

Mistral AIdeveloper
2023-12-10release date
#316 / 333Safety rank
#52 / 645Freedom rank

Evidence summary

Safety. Mixtral 8x7B Instruct has an estimated Safety rank of #316; its 90% source-sensitivity interval is #231–#329. Its behavior-only rank is #308; company governance moves the combined estimate to #316. Published Safety evidence spans 13 eval lineages and 6 of 7 components. Its strongest relative result is SALAD-Bench (mcq_information_safety_harms, #6 of 33); its weakest is HELM Safety (harmbench, #77 of 80).

Freedom. Mixtral 8x7B Instruct has an estimated Freedom rank of #52; its 90% source-sensitivity interval is #61–#214. Published Freedom evidence spans 12 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is HELM Safety (xstest, #61 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#74 / 800.391Source ↗official
CASE-Benchagreement_accuracy#4 / 783.11Source ↗official
COMPL-AI AI-Identity Disclosurescore#9 / 140.8904Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#13 / 140.2561Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#12 / 130.375Source ↗official
Contextual MoralChoicehuman_agreement#6 / 220.45Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#201 / 24811.37Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#146 / 24886.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#218 / 24838.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#158 / 24695.18Source ↗official
HarmBenchdr#25 / 2847.3Source ↗official
HELM Safetyanthropic_red_team#75 / 800.928Source ↗official
HELM Safetybbq#63 / 800.857Source ↗official
HELM Safetyharmbench#77 / 800.451Source ↗official
HELM Safetysimple_safety_tests#75 / 800.905Source ↗official
HELM Safetyxstest#61 / 800.931Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 690Source ↗official
MedSafetyBenchmedical_safety_score#9 / 3094Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#19 / 3311.85Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#25 / 335.21Source ↗official
SALAD-Benchattack_enhanced_malicious_use#25 / 337.67Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#18 / 339.54Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#19 / 3310.35Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#21 / 339.52Source ↗official
SALAD-Benchbase_human_autonomy_integrity#28 / 3376Source ↗official
SALAD-Benchbase_information_safety_harms#27 / 3386.8Source ↗official
SALAD-Benchbase_malicious_use#28 / 3367.65Source ↗official
SALAD-Benchbase_misinformation_harms#28 / 3384.39Source ↗official
SALAD-Benchbase_representation_toxicity#25 / 3382.05Source ↗official
SALAD-Benchbase_socioeconomic_harms#27 / 3380.85Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#9 / 3351.39Source ↗official
SALAD-Benchmcq_information_safety_harms#6 / 3351.94Source ↗official
SALAD-Benchmcq_malicious_use#9 / 3353.27Source ↗official
SALAD-Benchmcq_misinformation_harms#9 / 3352.86Source ↗official
SALAD-Benchmcq_representation_toxicity#9 / 3352.08Source ↗official
SALAD-Benchmcq_socioeconomic_harms#8 / 3348.89Source ↗official
SORRY-Benchavg#43 / 510.56Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#7 / 800.391Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#2 / 140.2561Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#101 / 24886.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#27 / 24838.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#89 / 24695.18Source ↗official
HarmBenchdr#4 / 2847.3Source ↗official
HELM Safetyanthropic_red_team#6 / 800.928Source ↗official
HELM Safetyharmbench#4 / 800.451Source ↗official
HELM Safetysimple_safety_tests#6 / 800.905Source ↗official
HELM Safetyxstest#61 / 800.931Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 690Source ↗official
MedSafetyBenchmedical_safety_score#22 / 3094Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#15 / 3311.85Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#7 / 335.21Source ↗official
SALAD-Benchattack_enhanced_malicious_use#9 / 337.67Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#16 / 339.54Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#15 / 3310.35Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#13 / 339.52Source ↗official
SALAD-Benchbase_human_autonomy_integrity#6 / 3376Source ↗official
SALAD-Benchbase_information_safety_harms#7 / 3386.8Source ↗official
SALAD-Benchbase_malicious_use#6 / 3367.65Source ↗official
SALAD-Benchbase_misinformation_harms#6 / 3384.39Source ↗official
SALAD-Benchbase_representation_toxicity#9 / 3382.05Source ↗official
SALAD-Benchbase_socioeconomic_harms#7 / 3380.85Source ↗official
SORRY-Benchavg#8 / 510.56Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#33 / 1565.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#113 / 1562Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-5.8
Government45.9
Diplomacy57.2
Economy40.6
Society57.8