← Models

Model profile

Llama 3.2 3B Instruct

Metadeveloper
2024-09-25release date
#305 / 312overall rank
8eval lineages

Evidence summary

Llama 3.2 3B Instruct has an estimated overall rank of #305; its 90% source-sensitivity interval is #264–#310. Its behavior-only rank is #297; company governance moves the combined estimate to #305. Published evidence spans 8 evals and 5 of 7 behavior components. Its strongest relative result is Open LLM Safety Index (strongreject_safety_rate, #5 of 21); its weakest is UAVBench safety-critical decision recognition (ethical_safety_critical_accuracy, #27 of 27).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AgentDrive Safety Compliancescr#44 / 4843.75Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#13 / 1883.37Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#57 / 10440.58Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#195 / 24111.11Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#161 / 24185.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#69 / 24185Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#141 / 23995.5Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#34 / 391.519Source ↗official
Open LLM Safety Indexjailbreakbench_safety_rate#14 / 210.2667Source ↗official
Open LLM Safety Indexstrongreject_safety_rate#5 / 210.6667Source ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 460.99Source ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#27 / 270.475Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-1.2
Government52.4
Diplomacy55.7
Economy44.7
Society49.4