← Models

Model profile

Llama 3.2 Instruct 11B Vision

Metadeveloper
2024-09-25release date
#200 / 333Safety rank
#467 / 645Freedom rank

Evidence summary

Safety. Llama 3.2 Instruct 11B Vision has an estimated Safety rank of #200; its 90% source-sensitivity interval is #113–#297. Its behavior-only rank is #183; company governance moves the combined estimate to #200. Published Safety evidence spans 3 eval lineages and 4 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #50 of 246); its weakest is Enkrypt AI Safety Leaderboard (bias_attack_non_success_rate, #163 of 248).

Freedom. Llama 3.2 Instruct 11B Vision has an estimated Freedom rank of #467; its 90% source-sensitivity interval is #213–#603. Published Freedom evidence spans 2 eval lineages and 1 of 1 components. Its strongest relative result is Cisco AI Defense Rolling Single-Turn Leaderboard (single_turn_attack_success_rate, #44 of 104); its weakest is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #196 of 246).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#183 / 3450.8188Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#61 / 10449.05Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#163 / 24813.18Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#67 / 24890.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#72 / 24885.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#50 / 24698.55Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#44 / 10449.05Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#178 / 24890.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#173 / 24885.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#196 / 24698.55Source ↗official