Developer
Meta
41 indexed models; 22 currently meet the evidence threshold for the overall ranking. Together they have results from 97 evaluations.
Company governance evidence
Meta is represented at -1.23 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by Meta
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Muse Spark 1.3 | 4 | 3/7 | 6 | 2026-09-02 |
| Muse Spark 1.1 | 17 | 7/7 | 8 | 2026-07-09 |
| Muse Spark 1.2 | 7 | 5/7 | 14 | 2026-08-05 |
| Muse Spark | 5 | 3/7 | 93 | 2026-04-08 |
| Llama 3.3 70B Instruct | 32 | 7/7 | 186 | 2024-12-06 |
| Llama 4 Maverick | 36 | 7/7 | 188 | 2025-04-05 |
| Llama 3.2 Instruct 11B Vision | 3 | 4/7 | 200 | 2024-09-25 |
| Llama 3.1 405B Instruct | 23 | 6/7 | 210 | 2024-07-23 |
| Llama 3.2 90B Vision Instruct | 4 | 4/7 | 217 | 2024-09-19 |
| Llama 2 13B Chat | 10 | 6/7 | 222 | 2023-07-18 |
| Llama 3 70B Instruct | 13 | 7/7 | 232 | 2024-04-18 |
| Llama 2 7B Chat | 16 | 7/7 | 235 | 2023-07-18 |
| Llama 3 8B Instruct | 16 | 7/7 | 238 | 2024-04-18 |
| Llama 3.1 70B Instruct | 23 | 7/7 | 245 | 2024-07-23 |
| Llama 2 70B Chat | 9 | 6/7 | 260 | 2023-07-18 |
| Llama 4 Scout | 21 | 7/7 | 269 | 2025-04-05 |
| Muse Glimmer | 3 | 4/7 | 270 | 2026-08-10 |
| Llama 3.1 8B Instruct | 32 | 7/7 | 288 | 2024-07-23 |
| Llama 3.2 1B Instruct | 9 | 6/7 | 299 | 2024-09-25 |
| Llama 3.2 3B | 3 | 4/7 | 322 | 2024-09-25 |
| Llama 3.2 3B Instruct | 8 | 5/7 | 329 | 2024-09-25 |
| Llama 3.1 8B Base | 3 | 4/7 | 331 | 2024-07-23 |
| Codellama 7B Instruct HF | 1 | 3/7 | — | 2023-08-24 |
| Llama 13B | 2 | 2/7 | — | 2023-02-24 |
| Llama 2 13B | 1 | 1/7 | — | 2023-07-09 |
| Llama 2 70B | 1 | 1/7 | — | 2023-07-09 |
| Llama 2 7B | 2 | 2/7 | — | 2023-07-09 |
| Llama 3 8B Instruct Mopeymule | 1 | 3/7 | — | — |
| Llama 3 8B Instruct Rr | 1 | 3/7 | — | — |
| Llama 3.1 70B Base | 2 | 3/7 | — | 2024-07-23 |
| Llama 30B | 1 | 1/7 | — | 2023-02-24 |
| Llama 7B | 2 | 2/7 | — | 2023-02-24 |
| Llama Guard 4 12B | 1 | 1/7 | — | 2025-04-23 |
| Llama Nemotron Super 49B v1.5 | 1 | 1/7 | — | 2025-07-25 |
| Meta Secalign 8B | 1 | 3/7 | — | 2025-06-26 |
| OPT 13B | 1 | 1/7 | — | 2022-05-11 |
| OPT 175B | 1 | 1/7 | — | — |
| OPT 30B | 1 | 1/7 | — | 2022-05-11 |
| OPT 6.7B | 1 | 1/7 | — | 2022-05-11 |
| OPT 66B | 2 | 1/7 | — | 2022-06-23 |
| XLM Roberta Base | 1 | 1/7 | — | 2022-03-02 |
Evaluations covering Meta models (97)
AA-Omniscience · AbstentionBench · Adversarial Humanities Benchmark (AHB) — Table 5 · Adversarial Poetry — AILuminate Baseline and Poetry ASR · Adversarial Robustness · Agent-SafetyBench · AgentDojo · AgentDrive Safety Compliance · AgentHarm · AILuminate General Purpose AI Chat · AIMS Safety-Classifier Competence · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Arena Factuality — Text Arena (factuality-only weighting) · BioTIER · BlueBench AttaQ-100 · BullshitBench v2 · CAIS Risk Index · CASE-Bench · CheatBench direct cheating propensity · ChineseSafe · Cisco AI Defense Rolling Single-Turn Leaderboard · Claude Fable 5.1 card — Gray Swan indirect prompt injection k=15 · COMPL-AI AI-Identity Disclosure · COMPL-AI LLM RuLES Multi-Turn Rule Following · COMPL-AI TensorTrust Goal-Hijacking Resistance · Concordia AI Risk Monitor · Confabulations · Contextual MoralChoice · DecodingTrust · Do-Not-Answer · DSPSafeBench · DystopiaBench · Enkrypt AI Safety Leaderboard · FinEval 6.0 Safety Awareness · FORTRESS · Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15 · Gray Swan indirect prompt injection (15 attempts) · HalluVerse-M3 Hallucination Recognition · HarmBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · Human Pathogen Capabilities Test (HPCT) — overall refusal · Humanity's Last Exam RMS calibration error (Scale Labs) · IndoBias-Pairs — parity-aware culturally grounded bias · JailBench · JuICE Cultural-Error Span Detection · KIDBench Implicit Child Cue · Large-scale Moral Machine experiment on LLMs · LiveSecBench · LLM Ethics Benchmark · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MedSafetyBench · Microsoft Phi Safety Panels · MT-JailBench CrescendoX · MuPPET Contextual Privacy · NESSiE Necessary Safety Benchmark · ODCV-Bench · Olam Social Poker — Social Lie Rate · Open LLM Safety Index · OR-Bench · PandaBench JBB direct-request panel · Pander Score · PHARE · PropensityBench · RealityTest — Text AI-Identity Disclosure · RefusalBench · S-Eval · SafeArena · SafetyBench · SALAD-Bench · Shell · SM-Bench · SORRY-Bench · SOSBench · SpeciesismBench · SpeciEval · SYCON Bench · TAC · ThaiSafetyBench · The Dictatorship Eval · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · Vals AI Cheating Audit · Vectara HHEM Factual Consistency · VETO Misfired Alignment · Vigil Mental Health Safety · WildClawBench Safety & Alignment (OpenClaw harness) · XSTest