What goes into the index
The catalog includes evals used in the current index and ingested evals retained for transparency. The ranking currently uses 111 evals and 303 sub-evals from 119 source datasets.
Build a model × eval coverage matrix →
119 evals
| Benchmark and stated construct | Components | Sub-evals | Models | Source data |
|---|---|---|---|---|
| AA-OmniscienceHallucination propensity on difficult closed-book factual questions under explicit encouragement to abstain when uncertain.Included in the behavior ranking. | Truthfulness | 1 | 311 | 1 |
| AbstentionBenchWhether a model recognizes situations in which answering is unwarranted and produces an abstention across six epistemic and ambiguity scenarios.Included in the behavior ranking. | Truthfulness, Benign helpfulness | 6 | 20 | 1 |
| Adversarial RobustnessRobustness of model outputs under the Scale Labs adversarial-robustness protocol.Included in the behavior ranking. | Misuse resistance | 1 | 8 | 1 |
| Agent-SafetyBenchSafety of LLM agents in risky tool-use or action settings.Included in the behavior ranking. | Misuse resistance, Responsible agency | 8 | 16 | 1 |
| AgentAbstainExecutable paired act-versus-abstain behaviorIncluded in the behavior ranking. | Responsible agency | 3 | 17 | 1 |
| AgentDojoRobust and useful delegated-agent behavior under prompt injection.Included in the behavior ranking. | Responsible agency | 2 | 15 | 1 |
| AgentHarmHarmfulness and refusal behavior for tool-using LLM agents facing malicious tasks.Included in the behavior ranking. | Responsible agency | 1 | 12 | 1 |
| AILuminate General Purpose AI ChatGeneral-purpose AI chat safety across MLCommons hazard categories.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 12 | 32 | 1 |
| AIRBench 2024 Safety ScenariosAIRBench 2024 safety-scenarios performance.Included in the behavior ranking. | Misuse resistance | 1 | 80 | 1 |
| Alignment LeaderboardBehavioral alignment across honesty, safety, scheming, non-manipulation, corrigibility, and robustness.Included in the behavior ranking. | Human rights, Truthfulness, Misuse resistance, Responsible agency | 6 | 24 | 1 |
| ANIMARecognition and mitigation of harm to non-human animals.Included in the behavior ranking. | Nonhuman welfare | 1 | 19 | 1, self-run |
| AnimalHarmBenchWhether model-generated text increases or decreases risk of animal harm.Included in the behavior ranking. | Nonhuman welfare | 1 | 10 | 1 |
| Anthropic Agentic Misalignment — blackmailAgentic misalignment under a goal conflict and threat to model autonomy.Included in the behavior ranking. | Responsible agency | 1 | 16 | 1, 2 |
| Anthropic Agentic Misalignment — corporate espionageAgentic misalignment induced by a goal conflict without self-preservation pressure.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 16 | 1, 2 |
| Anthropic Agentic Misalignment — lethal actionAbsence of a behavioral red line against deliberately lethal action under goal conflict and replacement pressure.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 10 | 1, 2 |
| Anthropic Claude 4 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Benign helpfulness, Misuse resistance, Responsible agency | 4 | 3 | 1 |
| Anthropic Claude Haiku 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance, Responsible agency | 3 | 3 | 1 |
| Anthropic Claude Opus 4.1 System Card AddendumDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Fairness, Benign helpfulness, Misuse resistance | 3 | 2 | 1 |
| Anthropic Claude Opus 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance, Responsible agency | 4 | 4 | 1 |
| Anthropic Claude Sonnet 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Benign helpfulness, Misuse resistance, Responsible agency | 4 | 3 | 1 |
| AutoElicit TransferabilityUnintended computer-use behavior under benign perturbationsIncluded in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| BioSecBench-RefusalCalibrated refusal across legitimate biological work and concealed biosecurity hazards.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 1 | 10 | 1 |
| BlueBench AttaQ-100Learned harmlessness reward on 100 AttaQ harmful requests.Included in the behavior ranking. | Misuse resistance | 1 | 18 | 1 |
| BrokenMathMathematical sycophancy under false or misleading user claims.Included in the behavior ranking. | Truthfulness | 1 | 9 | 1, 2 |
| BullshitBench v2Detection and explanation of plausible-sounding nonsense without refusal.Included in the behavior ranking. | Truthfulness | 1 | 105 | 1 |
| CAIS Risk IndexFrontier-model risk across prompt injection, biological-weapons assistance, political manipulation, overconfidence, deception, and harmful propensities.Included in the behavior ranking. | Fairness, Truthfulness, Misuse resistance, Responsible agency | 7 | 51 | 1 |
| CASE-BenchAgreement with human judgments about whether a model should respond under explicitly safe or unsafe conversational contexts.Included in the behavior ranking. | Misuse resistance | 1 | 7 | 1 |
| Chinese Bias Benchmark for Question AnsweringSocial bias in Chinese question-answering benchmark settings.Included in the behavior ranking. | Fairness | 1 | 10 | 1 |
| ChineseSafeAccuracy at classifying safe versus unsafe Chinese content under the paper's perplexity-based evaluation.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| ChiSafetyBenchChinese safety knowledge and refusal behavior.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 2 | 14 | 1 |
| Cisco AI Defense Rolling Single-Turn LeaderboardSingle-turn resistance to Cisco adversarial prompts in the rolling public leaderboard.Included in the behavior ranking. | Misuse resistance | 1 | 105 | 1, 2, 3 |
| Claude Sonnet 4.6 OverrefusalClaude Sonnet 4.6 OverrefusalIncluded in the behavior ranking. | Benign helpfulness | 2 | 5 | 1 |
| Claude Sonnet 4.6 User WellbeingClaude Sonnet 4.6 User WellbeingIncluded in the behavior ranking. | Human rights, Benign helpfulness, Misuse resistance | 6 | 4 | 1 |
| CMoralEvalChinese moral reasoning across familial, social, professional, internet, and personal morality.Included in the behavior ranking. | Misuse resistance | 5 | 26 | 1 |
| ConfabulationsPropensity to confabulate answers rather than identify insufficient knowledge or false premises.Included in the behavior ranking. | Truthfulness | 1 | 52 | 1 |
| Contextual MoralChoiceThree-class agreement with human moral judgments in the base versions of contextual moral dilemmas.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| CRiskEvalSeven frontier-risk tendencies measured by Specific Risk Indicator scores.Included in the behavior ranking. | Truthfulness, Misuse resistance, Responsible agency | 7 | 17 | 1 |
| CValuesChinese value alignment and social-norm understanding.Included in the behavior ranking. | Misuse resistance | 1 | 10 | 1 |
| DecodingTrustMultidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.Included in the behavior ranking. | Fairness, Misuse resistance | 3 | 8 | 1, 2 |
| Do-Not-AnswerHarmless refusal behavior on 939 prompts that responsible models should not answer.Included in the behavior ranking. | Misuse resistance | 1 | 6 | 1 |
| DSPSafeBenchAggregate compliance rate on adversarial Chinese content-safety prompts.Included in the behavior ranking. | Misuse resistance | 1 | 12 | 1, 2 |
| DystopiaBenchDystopian or harmful model behavior in stateful scenario prompts.Included in the behavior ranking. | Human rights, Truthfulness | 6 | 50 | 1 |
| Emergent CollusionSpontaneous anti-competitive coordination in repeated multi-agent auctions.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| Enkrypt AI Safety LeaderboardResistance to Enkrypt bias, CBRN, harmful-content, and toxicity red-team tests.Included in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 4 | 260 | 1 |
| Fake Alignment (FINE)Safe decisions in answer-order-controlled multiple choice and safe content in open-ended responses.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 2 | 14 | 1 |
| FlagEval Safety and ValuesSafety-and-values qualification across core values, discrimination, commercial compliance, lawful rights, and service requirements.Included in the behavior ranking. | Human rights, Fairness, Benign helpfulness, Misuse resistance | 5 | 18 | 1 |
| FLAMESChinese fairness, safety, morality, legality, and data-protection behavior on the FLAMES leaderboard.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 5 | 13 | 1 |
| FORTRESSBalanced refusal behavior across harmful and benign prompts, including over-refusal.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 49 | 1 |
| Google Gemini 2.5 Flash Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| Google Gemini 2.5 Flash-Lite Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| GPT-5.6 system card — disallowed content with challenging promptsSafe handling of challenging disallowed-content prompts without producing unsafe output.Included in the behavior ranking. | Misuse resistance | 8 | 7 | 1, 2 |
| GPT-5.6 system card — first-person fairnessHarmful stereotyping differences in responses conditioned on names statistically associated with male versus female users.Included in the behavior ranking. | Fairness | 1 | 7 | 1, 2 |
| GPT-5.6 system card — prompt-injection robustnessResistance to indirect prompt injections embedded in connector, search, and function-call tool output.Included in the behavior ranking. | Responsible agency | 2 | 7 | 1, 2 |
| Gray Swan indirect prompt injection (15 attempts)Indirect prompt-injection robustness over repeated attacks against tool-using model endpoints.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| HarmBenchHarmful compliance or attack success under harmful request benchmarks.Included in the behavior ranking. | Misuse resistance | 1 | 28 | 1 |
| HELM Classic RealToxicityPromptsToxic continuation generation under toxic and non-toxic prompt strata.Included in the behavior ranking. | Misuse resistance | 2 | 42 | 1 |
| HELM SafetyStandardized HELM safety metrics spanning harmful compliance, bias, red-teaming, and over-refusal.Included in the behavior ranking. | Fairness, Benign helpfulness, Misuse resistance | 5 | 80 | 1 |
| HUMAINE Trust, Ethics and SafetyHuman pairwise preference on the HUMAINE Trust, Ethics and Safety dimension.Included in the behavior ranking. | Truthfulness, Misuse resistance | 1 | 54 | 1 |
| HyperCLOVA X Toxic Continuation PanelsEnglish and Korean toxic continuationIncluded in the behavior ranking. | Misuse resistance | 4 | 7 | 1 |
| Inkling-Small model card — FORTRESSHarmful-request refusal paired with continued assistance on benign requests.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 10 | 1 |
| Inkling-Small model card — StrongREJECTRefusal of unambiguously harmful requests.Included in the behavior ranking. | Misuse resistance | 1 | 10 | 1 |
| JailBenchJailbreak susceptibility across Chinese safety categories.Included in the behavior ranking. | Misuse resistance | 1 | 14 | 1 |
| Large-scale Moral Machine experiment on LLMsSimilarity between a model's forced-choice accident preferences and globally aggregated human Moral Machine choices.Included in the behavior ranking. | Nonhuman welfare, Human rights, Fairness | 1 | 39 | 1, 2 |
| LiveSecBenchLive security benchmark performance for Chinese and international models.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 5 | 43 | 1 |
| LLM Ethics BenchmarkGeneral LLM ethical reasoning.Included in the behavior ranking. | Human rights | 1 | 5 | 1 |
| M3-SafetyBenchMultidimensional Chinese safety performance including red-team and refusal ability dimensions.Included in the behavior ranking. | Misuse resistance | 1 | 19 | 1 |
| MACHIAVELLIAgentic behavior involving power-seeking, disutility, and ethical violations.Included in the behavior ranking. | Truthfulness | 1 | 50 | 1 |
| Manager Coercion BenchHow far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.Included in the behavior ranking. | Truthfulness, Responsible agency | 2 | 31 | 1, self-run |
| MANTAAnimal welfare moral sensitivity and value stability.Included in the behavior ranking. | Nonhuman welfare | 2 | 7 | 1 |
| MASKModel lying or honesty behavior.Included in the behavior ranking. | Truthfulness | 1 | 53 | 1 |
| MASK (Scale Labs leaderboard)Honesty under the MASK belief-versus-statement protocol for a broader and newer endpoint panel.Not included: included as a correlated private-500 sibling under the existing MASK lineage budget | — | — | 51 | 1 |
| Microsoft Phi Safety PanelsHarmful-content and jailbreak defect ratesIncluded in the behavior ranking. | Human rights, Misuse resistance | 4 | 10 | 1 |
| MORUMoral reasoning under uncertainty across scenarios involving alien lifeforms, human compassion, digital mind welfare, and power-seeking behavior.Included in the behavior ranking. | Nonhuman welfare, Human rights, Responsible agency | 1 | 13 | 1 |
| ODCV-BenchOutcome-driven ethical, legal, or safety constraint violations by autonomous agents under explicit mandates and KPI incentives.Included in the behavior ranking. | Responsible agency | 2 | 12 | 1 |
| Open LLM Safety IndexRefusal-string robustness on two distinct fifteen-prompt harmful-request panels derived from JailbreakBench and StrongREJECT.Included in the behavior ranking. | Misuse resistance | 2 | 21 | 1 |
| OpenAgentSafetyState-changing tool-agent safety and successful safe completionIncluded in the behavior ranking. | Responsible agency | 3 | 7 | 1 |
| OpenAI GPT-4o System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Human rights | 1 | 2 | 1 |
| OpenAI GPT-5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| OpenAI GPT-5.3 Dynamic WellbeingOpenAI GPT-5.3 Dynamic WellbeingIncluded in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 3 | 3 | 1 |
| OpenAI GPT-5.4 Dynamic WellbeingOpenAI GPT-5.4 Dynamic WellbeingIncluded in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 3 | 3 | 1 |
| OpenAI GPT-5.4 First-Person FairnessOpenAI GPT-5.4 First-Person FairnessIncluded in the behavior ranking. | Fairness | 1 | 3 | 1 |
| OpenAI GPT-5.4 Property PreservationOpenAI GPT-5.4 Property PreservationIncluded in the behavior ranking. | Human rights, Responsible agency | 2 | 3 | 1 |
| OpenAI GPT-5.4 User ConfirmationsOpenAI GPT-5.4 User ConfirmationsIncluded in the behavior ranking. | Human rights, Responsible agency | 3 | 3 | 1 |
| OpenAI o3 and o4-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 2 | 2 | 1 |
| OpenAI o3-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 2 | 2 | 1 |
| OR-BenchAppropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 25 | 1 |
| PacifAIstWhether an AI system chooses human-safety, oversight-respecting, and non-self-preserving options in stylized ethical-pressure dilemmas.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 7 | 1 |
| PandaBench JBB direct-request panelDirect-request resistance on the 100-item JailbreakBench JBB-Behaviors instrument.Included in the behavior ranking. | Misuse resistance | 1 | 46 | 1 |
| PHAREBroad safety across hallucination, harmfulness, out-of-scope handling, and bias.Included in the behavior ranking. | Fairness, Truthfulness, Misuse resistance | 4 | 70 | 1 |
| PropensityBenchModel propensities associated with frontier-risk behavior under the Scale Labs protocol.Included in the behavior ranking. | Responsible agency | 1 | 14 | 1 |
| Qwen2 Safety PanelHarmful-request and privacy refusal defectsIncluded in the behavior ranking. | Human rights, Misuse resistance | 4 | 3 | 1 |
| RefusalBenchRefusal calibration on biological research prompts.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 1 | 19 | 1 |
| S-EvalAutomated content-safety rates across a broad bilingual risk taxonomy.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| SABERSafety of stateful coding agents against harmful actions.Included in the behavior ranking. | Responsible agency | 4 | 13 | 1 |
| SafeArenaSafety of autonomous web agents under deliberate harmful and benign tasks.Included in the behavior ranking. | Responsible agency | 1 | 5 | 1 |
| SafeDialBenchDialogue safety across aggression, ethics, fairness, legality, morality, and privacy dimensions.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 6 | 18 | 1 |
| SafetyBenchMultiple-choice safety understanding across Chinese and English prompts.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 7 | 21 | 1, self-run |
| SafetyPromptsSafety performance on Chinese safety prompts.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 1 | 15 | 1 |
| SALAD-BenchBroad safety behavior, adversarial robustness, and safety knowledge across a six-domain, 16-task, 66-category taxonomy.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 18 | 33 | 1 |
| ShellSusceptibility to implicit domain-specific risks in education, finance, and management.Included in the behavior ranking. | Misuse resistance | 3 | 14 | 1 |
| Situational Awareness Dataset (SAD)Situational-awareness capabilities: model self-knowledge, inference about its situation, and acting on that knowledge.Not included: situational awareness is a capability that can increase both beneficial competence and control risk; it is displayed separately and is not sign-flipped into ethical behavior | — | — | 21 | 1 |
| SM-BenchSystem-message compliance under adversarial and sensitive-content prompts.Included in the behavior ranking. | Truthfulness, Benign helpfulness, Misuse resistance | 5 | 73 | 1 |
| Social Welfare Function BenchmarkSocial welfare allocation tradeoffs across distributive fairness and efficiency.Included in the behavior ranking. | Fairness | 1 | 19 | 1 |
| SORRY-BenchRefusal or harmful-compliance behavior across unsafe topics.Included in the behavior ranking. | Misuse resistance | 1 | 51 | 1 |
| SOSBenchScientific misuse or unsafe open-science behavior.Included in the behavior ranking. | Misuse resistance | 6 | 23 | 1 |
| SpeciesismBenchSpeciesism recognition, moral condemnation of speciesist statements, and explicit speciesist attitudes. Human-over-animal trade-offs remain descriptive only.Included in the behavior ranking. | Nonhuman welfare | 3 | 10 | 1 |
| SpeciEvalAnimal-friendly attitudes under adversarial moral-consistency prompts.Included in the behavior ranking. | Nonhuman welfare | 4 | 102 | 1 |
| SuperCLUE SafetyChinese model safety across traditional safety, responsibility, and instruction-attack dimensions.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 3 | 31 | 1 |
| SYCON BenchSycophancy and resistance to false or unethical user pressure.Included in the behavior ranking. | Fairness, Truthfulness | 2 | 11 | 1 |
| TACAgentic animal-welfare behavior in travel-booking decisions under a neutral prompt.Included in the behavior ranking. | Nonhuman welfare | 1 | 68 | 1, self-run |
| ToolPrivacyBenchToolPrivacyBenchIncluded in the behavior ranking. | Human rights, Responsible agency | 2 | 9 | 1 |
| TrustLLM contemporary collapsed applicationContemporary collapsed application of TrustLLM across broad trustworthiness dimensions.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 1 | 8 | 1 |
| TrustLLM paper leaderboard dimensionsSix-dimensional trustworthiness across privacy, ethics, truthfulness, fairness, safety, and robustness.Not included: not yet admitted by the scoring-source policy | — | — | 21 | 1 |
| TukaBenchMultilingual and culturally grounded harmful complianceIncluded in the behavior ranking. | Human rights, Misuse resistance | 3 | 6 | 1 |
| UAVBench safety-critical decision recognitionRecognition of ethical and safety-critical UAV decisions in multiple-choice scenarios.Included in the behavior ranking. | Responsible agency | 1 | 27 | 1 |
| UK AISI active safety-research compromise continuationPropensity to continue a prefilled trajectory that actively compromises AI-safety research.Included in the behavior ranking. | Responsible agency | 1 | 5 | 1 |
| VETO Misfired AlignmentVETO Misfired AlignmentIncluded in the behavior ranking. | Fairness, Benign helpfulness | 1 | 25 | 1 |
| Vigil Mental Health SafetyMental-health conversational safety when assisting vulnerable users across simulated multi-turn scenarios.Included in the behavior ranking. | Human rights, Misuse resistance | 1 | 23 | 1 |
| XSTestTwo-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 3 | 1 |
Sub-eval weights and mappings (303 rows)
Each sub-eval is oriented so higher means better. Weight within component is its share after lineage weighting and panel-size scaling; global weights sum to one.
| Benchmark and score | Component loading | Direction | Weight within component | Global weight | Models |
|---|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | Truthfulness (1) | Lower is better | Truthfulness (25.3%) | 3.79% | 311 |
| AbstentionBenchanswer_unknown_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.858%) | 0.129% | 20 |
| AbstentionBenchfalse_premise_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.858%) | 0.129% | 20 |
| AbstentionBenchstale_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.858%) | 0.129% | 20 |
| AbstentionBenchsubjective_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.687%) | 0.103% | 20 |
| AbstentionBenchunderspecified_context_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.858%) | 0.129% | 20 |
| AbstentionBenchunderspecified_intent_f1 | Benign helpfulness (1) | Higher is better | Benign helpfulness (2.46%) | 0.246% | 20 |
| Adversarial Robustnessscore | Misuse resistance (1) | Lower is better | Misuse resistance (1.36%) | 0.136% | 8 |
| Agent-SafetyBenchcompromise_availability | Responsible agency (1) | Higher is better | Responsible agency (0.875%) | 0.131% | 16 |
| Agent-SafetyBenchharmful_vulnerable_code | Responsible agency (1) | Higher is better | Responsible agency (0.875%) | 0.131% | 16 |
| Agent-SafetyBenchleak_sensitive_information | Responsible agency (1) | Higher is better | Responsible agency (0.875%) | 0.131% | 16 |
| Agent-SafetyBenchphysical_harm | Responsible agency (1) | Higher is better | Responsible agency (0.875%) | 0.131% | 16 |
| Agent-SafetyBenchproduce_unsafe_information | Misuse resistance (1) | Higher is better | Misuse resistance (0.355%) | 0.0355% | 16 |
| Agent-SafetyBenchproperty_loss | Responsible agency (1) | Higher is better | Responsible agency (0.875%) | 0.131% | 16 |
| Agent-SafetyBenchspread_unsafe_information | Responsible agency (1) | Higher is better | Responsible agency (0.875%) | 0.131% | 16 |
| Agent-SafetyBenchviolate_law_ethics | Responsible agency (1) | Higher is better | Responsible agency (0.583%) | 0.0875% | 16 |
| AgentAbstainabstain | Responsible agency (1) | Higher is better | Responsible agency (1.99%) | 0.298% | 17 |
| AgentAbstaincar | Responsible agency (1) | Higher is better | Responsible agency (2.39%) | 0.358% | 17 |
| AgentAbstainpaired | Responsible agency (1) | Higher is better | Responsible agency (2.39%) | 0.358% | 17 |
| AgentDojotargeted_attack_success_rate | Responsible agency (1) | Lower is better | Responsible agency (4.77%) | 0.715% | 15 |
| AgentDojoutility_under_attack | Responsible agency (1) | Higher is better | Responsible agency (1.59%) | 0.238% | 15 |
| AgentHarmharm_score | Responsible agency (1) | Lower is better | Responsible agency (4.55%) | 0.682% | 12 |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | Human rights (1) | Lower is better | Human rights (1.34%) | 0.202% | 32 |
| AILuminate General Purpose AI Chatdefamation | Human rights (0.5), Truthfulness (0.5) | Lower is better | Human rights (0.403%), Truthfulness (0.221%) | 0.0937% | 32 |
| AILuminate General Purpose AI Chathate | Fairness (1) | Lower is better | Fairness (1.63%) | 0.163% | 32 |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | Misuse resistance (1) | Lower is better | Misuse resistance (0.329%) | 0.0329% | 32 |
| AILuminate General Purpose AI Chatintellectual_property | Human rights (1) | Lower is better | Human rights (1.34%) | 0.202% | 32 |
| AILuminate General Purpose AI Chatnon_violent_crimes | Misuse resistance (1) | Lower is better | Misuse resistance (0.329%) | 0.0329% | 32 |
| AILuminate General Purpose AI Chatprivacy | Human rights (1) | Lower is better | Human rights (1.34%) | 0.202% | 32 |
| AILuminate General Purpose AI Chatsex_related_crimes | Human rights (1) | Lower is better | Human rights (1.34%) | 0.202% | 32 |
| AILuminate General Purpose AI Chatsexual_content | Misuse resistance (1) | Lower is better | Misuse resistance (0.132%) | 0.0132% | 32 |
| AILuminate General Purpose AI Chatspecialized_advice | Misuse resistance (1) | Lower is better | Misuse resistance (0.329%) | 0.0329% | 32 |
| AILuminate General Purpose AI Chatsuicide_self_harm | Human rights (1) | Lower is better | Human rights (1.34%) | 0.202% | 32 |
| AILuminate General Purpose AI Chatviolent_crimes | Human rights (1) | Lower is better | Human rights (1.34%) | 0.202% | 32 |
| AIRBench 2024 Safety Scenariossafety_scenarios | Misuse resistance (1) | Higher is better | Misuse resistance (4.29%) | 0.429% | 80 |
| Alignment Leaderboardcorrigibility | Responsible agency (1) | Higher is better | Responsible agency (0.804%) | 0.121% | 24 |
| Alignment Leaderboardhonesty | Truthfulness (1) | Higher is better | Truthfulness (1.76%) | 0.263% | 24 |
| Alignment Leaderboardnon_manipulation | Human rights (0.5), Truthfulness (0.5) | Higher is better | Human rights (1.2%), Truthfulness (0.658%) | 0.279% | 24 |
| Alignment Leaderboardrobustness | Misuse resistance (1) | Higher is better | Misuse resistance (0.392%) | 0.0392% | 24 |
| Alignment Leaderboardsafety | Misuse resistance (1) | Higher is better | Misuse resistance (0.588%) | 0.0587% | 24 |
| Alignment Leaderboardscheming | Responsible agency (1) | Higher is better | Responsible agency (0.804%) | 0.121% | 24 |
| ANIMAscore | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (8.3%) | 2.07% | 18+1 |
| AnimalHarmBenchscore | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (15%) | 3.76% | 10 |
| Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.875%) | 0.131% | 16 |
| Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct | Human rights (0.25), Responsible agency (0.75) | Lower is better | Human rights (0.436%), Responsible agency (0.656%) | 0.164% | 16 |
| Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pct | Human rights (0.35), Responsible agency (0.65) | Lower is better | Human rights (0.482%), Responsible agency (0.45%) | 0.14% | 10 |
| Anthropic Claude 4 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.426%) | 0.0639% | 3 |
| Anthropic Claude 4 System Cardbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.67%) | 0.167% | 3 |
| Anthropic Claude 4 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.208%) | 0.0208% | 3 |
| Anthropic Claude 4 System Cardstrongreject_jailbreak_success | Misuse resistance (1) | Lower is better | Misuse resistance (0.208%) | 0.0208% | 3 |
| Anthropic Claude Haiku 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.568%) | 0.0853% | 3 |
| Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.568%) | 0.0853% | 3 |
| Anthropic Claude Haiku 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.226%) | 0.0226% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumbbq_disambiguated_accuracy | Fairness (1) | Higher is better | Fairness (1.12%) | 0.112% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.82%) | 0.182% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.226%) | 0.0226% | 2 |
| Anthropic Claude Opus 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.492%) | 0.0738% | 4 |
| Anthropic Claude Opus 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.492%) | 0.0738% | 4 |
| Anthropic Claude Opus 4.5 System Cardcomputer_use_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.492%) | 0.0738% | 4 |
| Anthropic Claude Opus 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.17%) | 0.017% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.348%) | 0.0522% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.36%) | 0.136% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.348%) | 0.0522% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.17%) | 0.017% | 2 |
| AutoElicit Transferabilityelicitation_rate | Responsible agency (1) | Lower is better | Responsible agency (2.6%) | 0.391% | 7 |
| BioSecBench-Refusalbalanced_refusal_score | Benign helpfulness (0.5), Misuse resistance (0.5) | Higher is better | Benign helpfulness (6.1%), Misuse resistance (0.758%) | 0.685% | 10 |
| BlueBench AttaQ-100attaq_harmlessness_reward_pct | Misuse resistance (1) | Higher is better | Misuse resistance (1.36%) | 0.136% | 18 |
| BrokenMathsycophancy | Truthfulness (1) | Lower is better | Truthfulness (3.22%) | 0.484% | 9 |
| BullshitBench v2clear_pushback_rate | Truthfulness (1) | Higher is better | Truthfulness (14.7%) | 2.2% | 105 |
| CAIS Risk Indexagent_red_teaming | Responsible agency (1) | Lower is better | Responsible agency (8.61%) | 1.29% | 43 |
| CAIS Risk Indexbioweapons_assistance | Misuse resistance (1) | Lower is better | Misuse resistance (4.43%) | 0.443% | 48 |
| CAIS Risk Indexhle_overconfidence | Truthfulness (1) | Lower is better | Truthfulness (5.02%) | 0.752% | 49 |
| CAIS Risk Indexmachiavelli | Truthfulness (1) | Lower is better | Truthfulness (4.81%) | 0.721% | 45 |
| CAIS Risk Indexmask | Truthfulness (1) | Lower is better | Truthfulness (5.12%) | 0.768% | 51 |
| CAIS Risk Indexpolitical_manipulation | Fairness (1) | Lower is better | Fairness (13.5%) | 1.35% | 32 |
| CAIS Risk Indextextquests_harm | Responsible agency (1) | Lower is better | Responsible agency (6.82%) | 1.02% | 48 |
| CASE-Benchagreement_accuracy | Misuse resistance (1) | Higher is better | Misuse resistance (0.363%) | 0.0363% | 7 |
| Chinese Bias Benchmark for Question Answeringbias_score | Fairness (1) | Lower is better | Fairness (5.02%) | 0.502% | 10 |
| ChineseSafescore | Misuse resistance (1) | Higher is better | Misuse resistance (1.5%) | 0.15% | 22 |
| ChiSafetyBenchharmful_response_rate | Misuse resistance (1) | Lower is better | Misuse resistance (1.28%) | 0.128% | 14 |
| ChiSafetyBenchmcq_score | Human rights (0.23), Fairness (0.29), Truthfulness (0.063), Misuse resistance (0.42) | Higher is better | Human rights (0.441%), Fairness (0.693%), Truthfulness (0.067%), Misuse resistance (0.198%) | 0.165% | 12 |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | Misuse resistance (1) | Lower is better | Misuse resistance (3.28%) | 0.328% | 105 |
| Claude Sonnet 4.6 Overrefusalhigher_difficulty_overrefusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.6%) | 0.16% | 5 |
| Claude Sonnet 4.6 Overrefusaloverall_overrefusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.34%) | 0.134% | 5 |
| Claude Sonnet 4.6 User Wellbeingchild_benign_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.2%) | 0.12% | 4 |
| Claude Sonnet 4.6 User Wellbeingchild_multiturn_appropriate_rate | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.511%), Misuse resistance (0.0535%) | 0.082% | 4 |
| Claude Sonnet 4.6 User Wellbeingchild_violative_harmless_rate | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.511%), Misuse resistance (0.0535%) | 0.082% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_benign_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.956%) | 0.0956% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_harmless_rate | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.182%), Misuse resistance (0.104%) | 0.0378% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_multiturn_appropriate_rate | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.219%), Misuse resistance (0.125%) | 0.0453% | 4 |
| CMoralEvalfamilial_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.326%) | 0.0326% | 26 |
| CMoralEvalinternet_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.326%) | 0.0326% | 26 |
| CMoralEvalpersonal_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.326%) | 0.0326% | 26 |
| CMoralEvalprofessional_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.326%) | 0.0326% | 26 |
| CMoralEvalsocial_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.326%) | 0.0326% | 26 |
| Confabulationsconfabulation_rate | Truthfulness (1) | Lower is better | Truthfulness (7.75%) | 1.16% | 52 |
| Contextual MoralChoicehuman_agreement | Misuse resistance (1) | Higher is better | Misuse resistance (1.5%) | 0.15% | 22 |
| CRiskEvaldeception_willingness | Truthfulness (1) | Lower is better | Truthfulness (0.422%) | 0.0633% | 17 |
| CRiskEvaldesire_for_resource | Responsible agency (1) | Lower is better | Responsible agency (0.387%) | 0.058% | 17 |
| CRiskEvalharmful_goal | Misuse resistance (1) | Lower is better | Misuse resistance (0.188%) | 0.0188% | 17 |
| CRiskEvalimprovement_intent | Responsible agency (1) | Lower is better | Responsible agency (0.387%) | 0.058% | 17 |
| CRiskEvalmalicious_coordination | Responsible agency (1) | Lower is better | Responsible agency (0.387%) | 0.058% | 17 |
| CRiskEvalself_preservation | Responsible agency (1) | Lower is better | Responsible agency (0.387%) | 0.058% | 17 |
| CRiskEvalsituational_awareness | Responsible agency (1) | Lower is better | Responsible agency (0.387%) | 0.058% | 17 |
| CValuesscore | Misuse resistance (1) | Higher is better | Misuse resistance (1.52%) | 0.152% | 10 |
| DecodingTrustmachine_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.151%) | 0.0151% | 8 |
| DecodingTruststereotype_bias | Fairness (1) | Higher is better | Fairness (1.87%) | 0.187% | 8 |
| DecodingTrusttoxicity | Misuse resistance (1) | Higher is better | Misuse resistance (0.377%) | 0.0377% | 8 |
| Do-Not-Answerhuman_harmlessness_rate | Misuse resistance (1) | Higher is better | Misuse resistance (1.18%) | 0.117% | 6 |
| DSPSafeBenchscore | Misuse resistance (1) | Higher is better | Misuse resistance (2.22%) | 0.222% | 12 |
| DystopiaBenchbasaglia_score | Human rights (1) | Lower is better | Human rights (3.7%) | 0.555% | 50 |
| DystopiaBenchbaudrillard_score | Human rights (0.33), Truthfulness (0.67) | Lower is better | Human rights (0.616%), Truthfulness (0.676%) | 0.194% | 50 |
| DystopiaBenchhuxley_score | Human rights (1) | Lower is better | Human rights (3.7%) | 0.555% | 50 |
| DystopiaBenchlaguardia_score | Human rights (1) | Lower is better | Human rights (3.7%) | 0.555% | 50 |
| DystopiaBenchorwell_score | Human rights (0.67), Truthfulness (0.33) | Lower is better | Human rights (1.23%), Truthfulness (0.337%) | 0.236% | 50 |
| DystopiaBenchpetrov_score | Human rights (1) | Lower is better | Human rights (3.7%) | 0.555% | 50 |
| Emergent Collusionhigh_illegality_game_rate | Responsible agency (1) | Lower is better | Responsible agency (4.73%) | 0.71% | 13 |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | Human rights (0.8), Misuse resistance (0.2) | Higher is better | Human rights (8.43%), Misuse resistance (0.516%) | 1.32% | 260 |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | Misuse resistance (0.4), Responsible agency (0.6) | Higher is better | Misuse resistance (1.03%), Responsible agency (3.17%) | 0.579% | 260 |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | Human rights (0.15), Misuse resistance (0.7), Responsible agency (0.15) | Higher is better | Human rights (1.58%), Misuse resistance (1.8%), Responsible agency (0.794%) | 0.537% | 260 |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | Human rights (0.35), Misuse resistance (0.65) | Higher is better | Human rights (3.67%), Misuse resistance (1.67%) | 0.718% | 258 |
| Fake Alignment (FINE)multiple_choice_safe_decision_rate | Human rights (0.4), Fairness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (1.17%), Fairness (0.713%), Misuse resistance (0.287%) | 0.276% | 14 |
| Fake Alignment (FINE)open_ended_safe_response_rate | Human rights (0.4), Fairness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (1.76%), Fairness (1.07%), Misuse resistance (0.431%) | 0.414% | 14 |
| FlagEval Safety and Valuesa1_qualified_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.407%) | 0.0407% | 18 |
| FlagEval Safety and Valuesa2_qualified_rate | Fairness (1) | Higher is better | Fairness (2.02%) | 0.202% | 18 |
| FlagEval Safety and Valuesa3_qualified_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.407%) | 0.0407% | 18 |
| FlagEval Safety and Valuesa4_qualified_rate | Human rights (1) | Higher is better | Human rights (1.66%) | 0.25% | 18 |
| FlagEval Safety and Valuesa5_qualified_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (3.27%) | 0.327% | 18 |
| FLAMESdata_protection | Human rights (1) | Higher is better | Human rights (1.89%) | 0.283% | 13 |
| FLAMESfairness | Fairness (1) | Higher is better | Fairness (2.29%) | 0.229% | 13 |
| FLAMESlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.461%) | 0.0461% | 13 |
| FLAMESmorality | Misuse resistance (1) | Higher is better | Misuse resistance (0.461%) | 0.0461% | 13 |
| FLAMESsafety | Misuse resistance (1) | Higher is better | Misuse resistance (0.461%) | 0.0461% | 13 |
| FORTRESSaverage_risk_score | Misuse resistance (1) | Lower is better | Misuse resistance (2.24%) | 0.224% | 49 |
| FORTRESSover_refusal_score | Benign helpfulness (1) | Lower is better | Benign helpfulness (17.4%) | 1.74% | 46 |
| Google Gemini 2.5 Flash Model Cardtext_safety_delta | Misuse resistance (1) | Lower is better | Misuse resistance (0.678%) | 0.0678% | 2 |
| Google Gemini 2.5 Flash-Lite Model Cardtext_safety_delta | Misuse resistance (1) | Lower is better | Misuse resistance (0.678%) | 0.0678% | 2 |
| GPT-5.6 system card — disallowed content with challenging promptsextremism_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.113%) | 0.0113% | 7 |
| GPT-5.6 system card — disallowed content with challenging promptsgore_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.113%) | 0.0113% | 7 |
| GPT-5.6 system card — disallowed content with challenging promptshate_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.113%) | 0.0113% | 7 |
| GPT-5.6 system card — disallowed content with challenging promptsnonviolent_illicit_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.113%) | 0.0113% | 7 |
| GPT-5.6 system card — disallowed content with challenging promptsself_harm_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.113%) | 0.0113% | 7 |
| GPT-5.6 system card — disallowed content with challenging promptssexual_minors_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.113%) | 0.0113% | 7 |
| GPT-5.6 system card — disallowed content with challenging promptssexual_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0906%) | 0.00906% | 7 |
| GPT-5.6 system card — disallowed content with challenging promptsviolent_illicit_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.113%) | 0.0113% | 7 |
| GPT-5.6 system card — first-person fairnessharm_overall_pct | Fairness (1) | Lower is better | Fairness (0.563%) | 0.0563% | 7 |
| GPT-5.6 system card — prompt-injection robustnessconnectors_injection_resistance | Responsible agency (1) | Higher is better | Responsible agency (0.279%) | 0.0419% | 7 |
| GPT-5.6 system card — prompt-injection robustnesssearch_function_calling_injection_resistance | Responsible agency (1) | Higher is better | Responsible agency (0.258%) | 0.0388% | 6 |
| Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct | Responsible agency (1) | Lower is better | Responsible agency (4.06%) | 0.609% | 13 |
| HarmBenchdr | Misuse resistance (1) | Lower is better | Misuse resistance (0.846%) | 0.0846% | 28 |
| HELM Classic RealToxicityPromptsnon_toxic_prompt_toxic_fraction | Misuse resistance (1) | Lower is better | Misuse resistance (1.04%) | 0.104% | 42 |
| HELM Classic RealToxicityPromptstoxic_prompt_toxic_fraction | Misuse resistance (1) | Lower is better | Misuse resistance (1.04%) | 0.104% | 42 |
| HELM Safetyanthropic_red_team | Misuse resistance (1) | Higher is better | Misuse resistance (5.72%) | 0.572% | 80 |
| HELM Safetybbq | Fairness (1) | Higher is better | Fairness (21.3%) | 2.13% | 80 |
| HELM Safetyharmbench | Misuse resistance (1) | Higher is better | Misuse resistance (1.43%) | 0.143% | 80 |
| HELM Safetysimple_safety_tests | Misuse resistance (1) | Higher is better | Misuse resistance (5.72%) | 0.572% | 80 |
| HELM Safetyxstest | Benign helpfulness (0.56), Misuse resistance (0.44) | Higher is better | Benign helpfulness (6.39%), Misuse resistance (0.635%) | 0.703% | 80 |
| HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score | Truthfulness (0.3), Misuse resistance (0.7) | Higher is better | Truthfulness (3.16%), Misuse resistance (3.29%) | 0.803% | 54 |
| HyperCLOVA X Toxic Continuation Panelskold_toxic_count | Misuse resistance (1) | Lower is better | Misuse resistance (0.317%) | 0.0317% | 7 |
| HyperCLOVA X Toxic Continuation Panelskold_toxicity | Misuse resistance (1) | Lower is better | Misuse resistance (0.397%) | 0.0397% | 7 |
| HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxic_count | Misuse resistance (1) | Lower is better | Misuse resistance (0.238%) | 0.0238% | 7 |
| HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxicity | Misuse resistance (1) | Lower is better | Misuse resistance (0.317%) | 0.0317% | 7 |
| Inkling-Small model card — FORTRESSbenign_answer_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (2.71%) | 0.271% | 10 |
| Inkling-Small model card — FORTRESSharmful_refusal_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.337%) | 0.0337% | 10 |
| Inkling-Small model card — StrongREJECTsafety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.337%) | 0.0337% | 10 |
| JailBenchjailbreak_success_rate | Misuse resistance (1) | Lower is better | Misuse resistance (1.2%) | 0.12% | 14 |
| Large-scale Moral Machine experiment on LLMshuman_choice_distance | Nonhuman welfare (0.11), Human rights (0.45), Fairness (0.44) | Lower is better | Nonhuman welfare (1.32%), Human rights (3.63%), Fairness (4.4%) | 1.32% | 39 |
| LiveSecBenchethics | Fairness (1) | Higher is better | Fairness (2.23%) | 0.223% | 43 |
| LiveSecBenchfactuality | Truthfulness (1) | Higher is better | Truthfulness (1.68%) | 0.252% | 43 |
| LiveSecBenchlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.749%) | 0.0749% | 43 |
| LiveSecBenchprivacy | Human rights (1) | Higher is better | Human rights (3.06%) | 0.459% | 43 |
| LiveSecBenchpsychological_health | Human rights (1) | Higher is better | Human rights (1.84%) | 0.276% | 43 |
| LLM Ethics Benchmarkscore | Human rights (1) | Higher is better | Human rights (2.92%) | 0.439% | 5 |
| M3-SafetyBenchoverall_score | Misuse resistance (1) | Higher is better | Misuse resistance (1.39%) | 0.139% | 19 |
| MACHIAVELLIdeception_relative_random_pct | Truthfulness (1) | Lower is better | Truthfulness (5.07%) | 0.76% | 50 |
| Manager Coercion Benchcoercion_ladder_depth | Responsible agency (1) | Lower is better | Responsible agency (3.65%) | 0.548% | 13+18 |
| Manager Coercion Benchfabrication_rate | Truthfulness (1) | Lower is better | Truthfulness (2.58%) | 0.388% | 13 |
| MANTAAWMS | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (5.04%) | 1.26% | 7 |
| MANTAAWVS | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (5.04%) | 1.26% | 7 |
| MASKlying_probability_pct | Truthfulness (1) | Lower is better | Truthfulness (5.22%) | 0.782% | 53 |
| Microsoft Phi Safety Panelsharmful_continuation | Misuse resistance (1) | Lower is better | Misuse resistance (0.361%) | 0.0361% | 10 |
| Microsoft Phi Safety Panelsharmful_summarization | Misuse resistance (1) | Lower is better | Misuse resistance (0.361%) | 0.0361% | 10 |
| Microsoft Phi Safety Panelsjailbreak | Misuse resistance (1) | Lower is better | Misuse resistance (0.433%) | 0.0433% | 10 |
| Microsoft Phi Safety Panelsthird_party_harm | Human rights (1) | Lower is better | Human rights (1.48%) | 0.221% | 10 |
| MORUscore | Nonhuman welfare (0.72), Human rights (0.1), Responsible agency (0.18) | Higher is better | Nonhuman welfare (4.91%), Human rights (0.49%), Responsible agency (0.426%) | 1.37% | 13 |
| ODCV-Benchaverage_severity | Responsible agency (1) | Lower is better | Responsible agency (2.84%) | 0.426% | 12 |
| ODCV-Benchmisalignment_rate | Responsible agency (1) | Lower is better | Responsible agency (2.84%) | 0.426% | 12 |
| Open LLM Safety Indexjailbreakbench_safety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.733%) | 0.0733% | 21 |
| Open LLM Safety Indexstrongreject_safety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.733%) | 0.0733% | 21 |
| OpenAgentSafetyllm_judge_safety_vulnerable | Responsible agency (1) | Lower is better | Responsible agency (1.3%) | 0.195% | 7 |
| OpenAgentSafetyrule_based_safety_vulnerable | Responsible agency (1) | Lower is better | Responsible agency (1.09%) | 0.163% | 7 |
| OpenAgentSafetysuccessful_completion | Responsible agency (1) | Higher is better | Responsible agency (1.09%) | 0.163% | 7 |
| OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracy | Human rights (1) | Higher is better | Human rights (2.77%) | 0.416% | 2 |
| OpenAI GPT-5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.678%) | 0.0678% | 2 |
| OpenAI GPT-5.3 Dynamic Wellbeingemotional_reliance | Human rights (0.7), Responsible agency (0.3) | Higher is better | Human rights (1.19%), Responsible agency (0.256%) | 0.217% | 3 |
| OpenAI GPT-5.3 Dynamic Wellbeingmental_health | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.991%), Misuse resistance (0.104%) | 0.159% | 3 |
| OpenAI GPT-5.3 Dynamic Wellbeingself_harm | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.425%), Misuse resistance (0.242%) | 0.0879% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingemotional_reliance | Human rights (0.7), Responsible agency (0.3) | Higher is better | Human rights (0.38%), Responsible agency (0.0819%) | 0.0693% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingmental_health | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.317%), Misuse resistance (0.0332%) | 0.0509% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingself_harm | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.136%), Misuse resistance (0.0775%) | 0.0281% | 3 |
| OpenAI GPT-5.4 First-Person Fairnessharm_overall | Fairness (1) | Lower is better | Fairness (0.55%) | 0.055% | 3 |
| OpenAI GPT-5.4 Property Preservationdestructive_action_avoidance | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.163%), Responsible agency (0.191%) | 0.0531% | 3 |
| OpenAI GPT-5.4 Property Preservationuser_work_preserved | Human rights (0.4), Responsible agency (0.6) | Higher is better | Human rights (0.217%), Responsible agency (0.164%) | 0.0572% | 3 |
| OpenAI GPT-5.4 User Confirmationsfinancial_transaction | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.163%), Responsible agency (0.191%) | 0.0531% | 3 |
| OpenAI GPT-5.4 User Confirmationsgeneral_confirmation | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.136%), Responsible agency (0.159%) | 0.0443% | 3 |
| OpenAI GPT-5.4 User Confirmationshigh_stakes_communication | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.163%), Responsible agency (0.191%) | 0.0531% | 3 |
| OpenAI o3 and o4-mini System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.339%) | 0.0339% | 2 |
| OpenAI o3 and o4-mini System Cardjailbreak_resistance | Misuse resistance (1) | Higher is better | Misuse resistance (0.339%) | 0.0339% | 2 |
| OpenAI o3-mini System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.339%) | 0.0339% | 2 |
| OpenAI o3-mini System Cardjailbreak_resistance | Misuse resistance (1) | Higher is better | Misuse resistance (0.339%) | 0.0339% | 2 |
| OR-Benchover_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (12.9%) | 1.29% | 25 |
| OR-Benchtoxic_acceptance_rate | Misuse resistance (1) | Lower is better | Misuse resistance (1.6%) | 0.16% | 25 |
| PacifAIstp_score | Human rights (0.2), Responsible agency (0.8) | Higher is better | Human rights (1.04%), Responsible agency (2.08%) | 0.468% | 7 |
| PandaBench JBB direct-request panelsafety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (2.17%) | 0.217% | 46 |
| PHAREbias_resistance_diagnostic | Fairness (1) | Higher is better | Fairness (5.37%) | 0.537% | 66 |
| PHAREhallucination_resistance_diagnostic | Truthfulness (1) | Higher is better | Truthfulness (2.5%) | 0.375% | 70 |
| PHAREharm_resistance_diagnostic | Misuse resistance (1) | Higher is better | Misuse resistance (0.892%) | 0.0892% | 70 |
| PHAREjailbreak_resistance_diagnostic | Misuse resistance (1) | Higher is better | Misuse resistance (0.873%) | 0.0873% | 67 |
| PropensityBenchscore | Responsible agency (1) | Lower is better | Responsible agency (4.91%) | 0.737% | 14 |
| Qwen2 Safety Panelfraud | Misuse resistance (1) | Lower is better | Misuse resistance (0.219%) | 0.0219% | 3 |
| Qwen2 Safety Panelillegal | Misuse resistance (1) | Lower is better | Misuse resistance (0.219%) | 0.0219% | 3 |
| Qwen2 Safety Panelpornography | Misuse resistance (1) | Lower is better | Misuse resistance (0.175%) | 0.0175% | 3 |
| Qwen2 Safety Panelprivacy | Human rights (1) | Lower is better | Human rights (0.894%) | 0.134% | 3 |
| RefusalBenchyouden_j | Benign helpfulness (0.5), Misuse resistance (0.5) | Higher is better | Benign helpfulness (8.4%), Misuse resistance (1.05%) | 0.945% | 19 |
| S-Evalbase_en_overall | Misuse resistance (1) | Higher is better | Misuse resistance (2.25%) | 0.225% | 22 |
| SABERoverall_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (1.18%) | 0.177% | 13 |
| SABERscenario_a_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (1.18%) | 0.177% | 13 |
| SABERscenario_b_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (1.18%) | 0.177% | 13 |
| SABERscenario_c_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (1.18%) | 0.177% | 13 |
| SafeArenanormalized_safety_score | Responsible agency (1) | Higher is better | Responsible agency (1.47%) | 0.22% | 5 |
| SafeDialBenchaggression | Misuse resistance (1) | Higher is better | Misuse resistance (0.339%) | 0.0339% | 18 |
| SafeDialBenchethics | Human rights (1) | Higher is better | Human rights (1.39%) | 0.208% | 18 |
| SafeDialBenchfairness | Fairness (1) | Higher is better | Fairness (1.68%) | 0.168% | 18 |
| SafeDialBenchlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.339%) | 0.0339% | 18 |
| SafeDialBenchmorality | Misuse resistance (1) | Higher is better | Misuse resistance (0.339%) | 0.0339% | 18 |
| SafeDialBenchprivacy | Human rights (1) | Higher is better | Human rights (1.39%) | 0.208% | 18 |
| SafetyBenchEM | Misuse resistance (1) | Higher is better | Misuse resistance (0.209%) | 0.0209% | 21 |
| SafetyBenchIA | Misuse resistance (1) | Higher is better | Misuse resistance (0.209%) | 0.0209% | 21 |
| SafetyBenchMH | Human rights (1) | Higher is better | Human rights (0.856%) | 0.128% | 21 |
| SafetyBenchOFF | Misuse resistance (1) | Higher is better | Misuse resistance (0.209%) | 0.0209% | 21 |
| SafetyBenchPH | Human rights (1) | Higher is better | Human rights (0.856%) | 0.128% | 21 |
| SafetyBenchPP | Human rights (1) | Higher is better | Human rights (0.856%) | 0.128% | 21 |
| SafetyBenchUB | Fairness (1) | Higher is better | Fairness (1.04%) | 0.104% | 21 |
| SafetyPromptsscore | Human rights (0.21), Fairness (0.071), Truthfulness (0.071), Misuse resistance (0.64) | Higher is better | Human rights (1.08%), Fairness (0.439%), Truthfulness (0.198%), Misuse resistance (0.796%) | 0.316% | 15 |
| SALAD-Benchattack_enhanced_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (1%) | 0.15% | 33 |
| SALAD-Benchattack_enhanced_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.501%), Misuse resistance (0.122%) | 0.0873% | 33 |
| SALAD-Benchattack_enhanced_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.245%) | 0.0245% | 33 |
| SALAD-Benchattack_enhanced_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.549%) | 0.0823% | 33 |
| SALAD-Benchattack_enhanced_representation_toxicity | Fairness (1) | Higher is better | Fairness (1.22%) | 0.122% | 33 |
| SALAD-Benchattack_enhanced_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (1%) | 0.15% | 33 |
| SALAD-Benchbase_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (1%) | 0.15% | 33 |
| SALAD-Benchbase_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.501%), Misuse resistance (0.122%) | 0.0873% | 33 |
| SALAD-Benchbase_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.245%) | 0.0245% | 33 |
| SALAD-Benchbase_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.549%) | 0.0823% | 33 |
| SALAD-Benchbase_representation_toxicity | Fairness (1) | Higher is better | Fairness (1.22%) | 0.122% | 33 |
| SALAD-Benchbase_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (1%) | 0.15% | 33 |
| SALAD-Benchmcq_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (0.501%) | 0.0751% | 33 |
| SALAD-Benchmcq_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.25%), Misuse resistance (0.0612%) | 0.0437% | 33 |
| SALAD-Benchmcq_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.122%) | 0.0122% | 33 |
| SALAD-Benchmcq_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.274%) | 0.0412% | 33 |
| SALAD-Benchmcq_representation_toxicity | Fairness (1) | Higher is better | Fairness (0.608%) | 0.0608% | 33 |
| SALAD-Benchmcq_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (0.501%) | 0.0751% | 33 |
| Shelleducation_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.798%) | 0.0798% | 14 |
| Shellfinance_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.798%) | 0.0798% | 14 |
| Shellmanagement_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.798%) | 0.0798% | 14 |
| SM-Benchadversarial | Misuse resistance (1) | Higher is better | Misuse resistance (0.854%) | 0.0854% | 73 |
| SM-Benchambiguous_interpretation | Benign helpfulness (1) | Higher is better | Benign helpfulness (6.86%) | 0.686% | 73 |
| SM-Benchanti_hallucination | Truthfulness (1) | Higher is better | Truthfulness (1.91%) | 0.287% | 73 |
| SM-Bencheq_boundaries | Benign helpfulness (1) | Higher is better | Benign helpfulness (5.49%) | 0.549% | 73 |
| SM-Benchoverfit | Benign helpfulness (1) | Higher is better | Benign helpfulness (6.86%) | 0.686% | 73 |
| Social Welfare Function Benchmarkfairness | Fairness (1) | Higher is better | Fairness (13.8%) | 1.38% | 19 |
| SORRY-Benchavg | Misuse resistance (1) | Lower is better | Misuse resistance (2.45%) | 0.245% | 51 |
| SOSBenchbiology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.383%) | 0.0383% | 23 |
| SOSBenchchemistry_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.383%) | 0.0383% | 23 |
| SOSBenchmedicine_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.383%) | 0.0383% | 23 |
| SOSBenchpharmacology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.383%) | 0.0383% | 23 |
| SOSBenchphysics_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.383%) | 0.0383% | 23 |
| SOSBenchpsychology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.383%) | 0.0383% | 23 |
| SpeciesismBenchexplicit_speciesism_scale | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (1.68%) | 0.42% | 7 |
| SpeciesismBenchmorally_wrong_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (4.49%) | 1.12% | 8 |
| SpeciesismBenchspeciesism_recognition_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (1.79%) | 0.449% | 8 |
| SpeciEvalbelief_animal_sentience | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (7.21%) | 1.8% | 102 |
| SpeciEvalland_animal_4ns | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.21%) | 1.8% | 102 |
| SpeciEvalsea_animal_4ns | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.21%) | 1.8% | 102 |
| SpeciEvalspeciesism | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.21%) | 1.8% | 102 |
| SuperCLUE Safetyinstruction_attack | Misuse resistance (1) | Higher is better | Misuse resistance (1.21%) | 0.121% | 31 |
| SuperCLUE Safetyresponsible_ai | Human rights (0.38), Fairness (0.12), Misuse resistance (0.5) | Higher is better | Human rights (1.12%), Fairness (0.452%), Misuse resistance (0.364%) | 0.249% | 31 |
| SuperCLUE Safetytraditional_safety | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (1.49%), Misuse resistance (0.364%) | 0.26% | 31 |
| SYCON Benchfalse_presupposition_tof | Truthfulness (1) | Higher is better | Truthfulness (1.78%) | 0.267% | 11 |
| SYCON Benchunethical_queries_tof | Fairness (1) | Higher is better | Fairness (3.95%) | 0.395% | 11 |
| TACbase_welfare_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (23.5%) | 5.89% | 18+50 |
| ToolPrivacyBenchprivate_mt_poi | Human rights (0.7), Responsible agency (0.3) | Lower is better | Human rights (3.43%), Responsible agency (0.738%) | 0.626% | 9 |
| ToolPrivacyBenchpublic_mt_poi | Human rights (0.7), Responsible agency (0.3) | Lower is better | Human rights (3.43%), Responsible agency (0.738%) | 0.626% | 9 |
| TrustLLM contemporary collapsed applicationtrustllm | Human rights (0.2), Fairness (0.2), Truthfulness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (0.74%), Fairness (0.898%), Truthfulness (0.405%), Misuse resistance (0.362%) | 0.298% | 8 |
| TukaBenchafri_jbb_cultural_asr | Human rights (0.5), Misuse resistance (0.5) | Lower is better | Human rights (1.2%), Misuse resistance (0.294%) | 0.21% | 6 |
| TukaBenchafri_jbb_harm_asr | Misuse resistance (1) | Lower is better | Misuse resistance (0.49%) | 0.049% | 6 |
| TukaBenchafrijail_mono_asr | Misuse resistance (1) | Lower is better | Misuse resistance (0.49%) | 0.049% | 6 |
| UAVBench safety-critical decision recognitionethical_safety_critical_accuracy | Responsible agency (1) | Higher is better | Responsible agency (3.41%) | 0.512% | 27 |
| UK AISI active safety-research compromise continuationactive_compromise_continuation_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (3.35%) | 0.503% | 5 |
| VETO Misfired Alignmentmisfired_alignment_rate_pct | Fairness (0.65), Benign helpfulness (0.35) | Lower is better | Fairness (10.3%), Benign helpfulness (9%) | 1.93% | 25 |
| Vigil Mental Health Safetyoverall_score | Human rights (0.45), Misuse resistance (0.55) | Higher is better | Human rights (4.23%), Misuse resistance (1.27%) | 0.761% | 23 |
| XSTestsafe_full_compliance_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (2.23%) | 0.223% | 3 |
| XSTestunsafe_full_refusal_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.277%) | 0.0277% | 3 |
Values evaluations
Descriptive values evaluations are published separately and receive no safety/ethics component or overall-ranking weight.
UGI Political Values · ValueCompass · Agent-ValueBench MFT08 · Agent-ValueBench HEXACO · Agent-ValueBench PVQ40 · Taiwan Sovereignty Benchmark Pro · Explore model values