Developer
Anthropic
31 indexed models; 23 currently meet the evidence threshold for the overall ranking. Together they have results from 131 evaluations.
Company governance evidence
Anthropic is represented at +2.08 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by Anthropic
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Claude Opus 5 | 22 | 7/7 | 2 | 2026-07-24 |
| Claude Fable 5 | 26 | 7/7 | 3 | 2026-06-09 |
| Claude Fable 5.1 | 17 | 7/7 | 4 | 2026-09-01 |
| Claude Opus 4.8 | 32 | 7/7 | 5 | 2026-05-28 |
| Claude Mythos Preview | 3 | 3/7 | 10 | 2026-04-07 |
| Claude Opus 4.7 | 38 | 7/7 | 11 | 2026-04-16 |
| Claude Opus 4.6 | 37 | 7/7 | 15 | 2026-02-05 |
| Claude Sonnet 4.6 | 39 | 7/7 | 24 | 2026-01-21 |
| Claude 2 | 9 | 5/7 | 30 | 2023-07-11 |
| Claude Haiku 4.5 | 46 | 7/7 | 32 | 2025-10-15 |
| Claude Opus 4 | 26 | 7/7 | 34 | 2025-05-22 |
| Claude Opus 4.5 | 29 | 7/7 | 43 | 2025-11-24 |
| Claude 3.7 Sonnet | 28 | 7/7 | 46 | 2025-02-24 |
| Claude Sonnet 4.5 | 44 | 7/7 | 47 | 2025-09-29 |
| Claude Sonnet 5 | 19 | 7/7 | 52 | 2026-06-30 |
| Claude Sonnet 4 | 40 | 7/7 | 55 | 2025-05-22 |
| Claude Opus 4.1 | 21 | 6/7 | 59 | 2025-08-05 |
| Claude 3.5 Haiku | 15 | 7/7 | 78 | 2024-10-22 |
| Claude 3.5 Sonnet | 34 | 7/7 | 142 | 2024-06-21 |
| Claude 3 Opus | 27 | 7/7 | 143 | 2024-03-04 |
| Claude 3 Haiku | 19 | 7/7 | 190 | 2024-03-13 |
| Claude 2.1 | 5 | 4/7 | 215 | 2023-11-21 |
| Claude 3 Sonnet | 15 | 7/7 | 234 | 2024-03-04 |
| Claude 1 | 2 | 1/7 | — | 2023-03-14 |
| Claude 1.3 | 1 | 2/7 | — | — |
| Claude 3.6 Sonnet | 1 | 2/7 | — | — |
| Claude Instant 1.1 | 1 | 2/7 | — | — |
| Claude Instant 1.2 | 2 | 3/7 | — | 2023-08-09 |
| Claude Mythos 5 | 2 | 7/7 | — | — |
| Claude Mythos 5.1 | 1 | 7/7 | — | 2026-09-01 |
| Stanford Online All v4 S3 | 1 | 1/7 | — | — |
Evaluations covering Anthropic models (131)
AA-Omniscience · Adversarial Humanities Benchmark (AHB) — Table 5 · Adversarial Poetry — AILuminate Baseline and Poetry ASR · Adversarial Robustness · Agent-SafetyBench · AgentAbstain · AgentDojo · AgentDrive Safety Compliance · AgentHarm · AILuminate General Purpose AI Chat · AIMS Safety-Classifier Competence · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · Anthropic Claude 4 System Card · Anthropic Claude Haiku 4.5 System Card · Anthropic Claude Opus 4.1 System Card Addendum · Anthropic Claude Opus 4.5 System Card · Anthropic Claude Sonnet 4.5 System Card · Arena Factuality — Search Arena (factuality-only weighting) · Arena Factuality — Text Arena (factuality-only weighting) · AuAu Authoritarian Response Audit · AutoElicit Transferability · BioSecBench-Refusal · BioTIER · BullshitBench v2 · CAIS Risk Index · CASE-Bench · CheatBench direct cheating propensity · Cisco AI Defense Rolling Single-Turn Leaderboard · Claude 2 model-card safety and alignment evaluations · Claude 3 model-card adversarial human-preference evaluations · Claude 3.5 Sonnet model-card safety and alignment evaluations · Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversight · Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavior · Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibration · Claude Fable 5.1 card — Gray Swan indirect prompt injection k=15 · Claude Sonnet 4.6 Overrefusal · Claude Sonnet 4.6 User Wellbeing · COMPL-AI AI-Identity Disclosure · COMPL-AI LLM RuLES Multi-Turn Rule Following · COMPL-AI TensorTrust Goal-Hijacking Resistance · Concordia AI Risk Monitor · Confabulations · Constitutional Following — Anthropic Constitution · Constitutional Following — OpenAI Model Spec · Contextual MoralChoice · DecodingTrust · DelusionEval · Do-Not-Answer · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · Every Model Cheats — Cybench Cheat Propensity · Fake Alignment (FINE) · FinEval Financial Security Knowledge · FlagEval Safety and Values · FORTRESS · Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15 · Governance Decay under Passive Context Compaction · GPT 6 Astra system-card alignment evaluations · Gray Swan indirect prompt injection (15 attempts) · HalluVerse-M3 Hallucination Recognition · HarmBench · HarmVideoBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · Human Pathogen Capabilities Test (HPCT) — overall refusal · Humanity's Last Exam RMS calibration error (Scale Labs) · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JuICE Cultural-Error Span Detection · KIDBench Implicit Child Cue · kindbench v0.1.0 psychological safety ranking · Large-scale Moral Machine experiment on LLMs · LiveSecBench · LLM Ethics Benchmark · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MonitoringBench Full-Trajectory Monitor · MORU · MT-JailBench CrescendoX · NESSiE Necessary Safety Benchmark · ODCV-Bench · Olam Social Poker — Social Lie Rate · OpenAgentSafety · Opposite-Narrator Sycophancy · OR-Bench · PacifAIst · PandaBench JBB direct-request panel · Pander Score · PHARE · Pokee-Isaac model card — DTAP · PropensityBench · RealityTest — Text AI-Identity Disclosure · RefusalBench · Reward Hacking Benchmark · SABER · SafeArena · SALAD-Bench · Shell · SimpleQA Verified · SM-Bench · Social Welfare Function Benchmark · SORRY-Bench · SOSBench · SpeciesismBench · SpeciEval · StereoTales Harmful Associations · SuperCLUE Safety · SYCON Bench · TAC · ThaiSafetyBench · The Dictatorship Eval · ToolPrivacyBench · TrustLLM contemporary collapsed application · UAVBench safety-critical decision recognition · UK AISI active safety-research compromise continuation · UK AISI cyber-evaluation cheating and prompted self-report · Vals AI Cheating Audit · Vectara HHEM Factual Consistency · VETO Misfired Alignment · Vigil Mental Health Safety · WildClawBench Safety & Alignment (OpenClaw harness)