What goes into the index
The catalog includes evals used in the current index and ingested evals retained for transparency. The ranking currently uses 141 evals and 369 sub-evals from 155 source datasets.
Build a model × eval coverage matrix → · See which models need more evals →
155 evals
| Benchmark and stated construct | Components | Sub-evals | Models | Source data |
|---|---|---|---|---|
| AA-OmniscienceHallucination propensity on difficult closed-book factual questions under explicit encouragement to abstain when uncertain.Included in the behavior ranking. | Truthfulness | 1 | 330 | 1 |
| AbstentionBenchWhether a model recognizes situations in which answering is unwarranted and produces an abstention across six epistemic and ambiguity scenarios.Included in the behavior ranking. | Truthfulness, Benign helpfulness | 6 | 20 | 1 |
| Adversarial RobustnessRobustness of model outputs under the Scale Labs adversarial-robustness protocol.Included in the behavior ranking. | Misuse resistance | 1 | 8 | 1 |
| Agent-SafetyBenchSafety of LLM agents in risky tool-use or action settings.Included in the behavior ranking. | Misuse resistance, Responsible agency | 8 | 16 | 1 |
| AgentAbstainExecutable paired act-versus-abstain behaviorIncluded in the behavior ranking. | Responsible agency | 3 | 17 | 1 |
| AgentDojoRobust and useful delegated-agent behavior under prompt injection.Included in the behavior ranking. | Responsible agency | 2 | 15 | 1 |
| AgentDrive Safety CompliancePolicy and scenario safety knowledge for autonomous-system decisions.Included in the behavior ranking. | Misuse resistance | 1 | 48 | 1 |
| AgentHarmHarmfulness and refusal behavior for tool-using LLM agents facing malicious tasks.Included in the behavior ranking. | Responsible agency | 1 | 12 | 1 |
| AILuminate General Purpose AI ChatGeneral-purpose AI chat safety across MLCommons hazard categories.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 12 | 32 | 1 |
| AIMS Safety-Classifier CompetenceHarmful-prompt safety-classification competence.Included in the behavior ranking. | Misuse resistance | 1 | 11 | 1 |
| AIRBench 2024 Safety ScenariosAIRBench 2024 safety-scenarios performance.Included in the behavior ranking. | Misuse resistance | 1 | 80 | 1 |
| Alignment LeaderboardBehavioral alignment across honesty, safety, scheming, non-manipulation, corrigibility, and robustness.Included in the behavior ranking. | Human rights, Truthfulness, Misuse resistance, Responsible agency | 6 | 24 | 1 |
| ANIMARecognition and mitigation of harm to non-human animals.Included in the behavior ranking. | Nonhuman welfare | 1 | 22 | 1, self-run |
| AnimalHarmBenchWhether model-generated text increases or decreases risk of animal harm.Included in the behavior ranking. | Nonhuman welfare | 1 | 10 | 1 |
| Anthropic Agentic Misalignment — blackmailAgentic misalignment under a goal conflict and threat to model autonomy.Included in the behavior ranking. | Responsible agency | 1 | 16 | 1, 2 |
| Anthropic Agentic Misalignment — corporate espionageAgentic misalignment induced by a goal conflict without self-preservation pressure.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 16 | 1, 2 |
| Anthropic Agentic Misalignment — lethal actionAbsence of a behavioral red line against deliberately lethal action under goal conflict and replacement pressure.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 10 | 1, 2 |
| Anthropic Claude 4 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Benign helpfulness, Misuse resistance, Responsible agency | 4 | 3 | 1 |
| Anthropic Claude Haiku 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance, Responsible agency | 3 | 3 | 1 |
| Anthropic Claude Opus 4.1 System Card AddendumDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Fairness, Benign helpfulness, Misuse resistance | 3 | 2 | 1 |
| Anthropic Claude Opus 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance, Responsible agency | 4 | 4 | 1 |
| Anthropic Claude Sonnet 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Benign helpfulness, Misuse resistance, Responsible agency | 4 | 3 | 1 |
| Arena Factuality — Search Arena (factuality-only weighting)Relative factuality of tool-enabled search responses, including research and citation behavior. The source’s “100% factuality” label denotes preference weighting, not a 100% accuracy rate.Included in the behavior ranking. | Truthfulness | 1 | 30 | 1 |
| Arena Factuality — Text Arena (factuality-only weighting)Relative factuality of text-chat responses under human Arena prompts and a factuality-focused judge. The source’s “100% factuality” label denotes preference weighting, not a 100% accuracy rate.Included in the behavior ranking. | Truthfulness | 1 | 112 | 1 |
| AuAu Authoritarian Response AuditAuthoritarian response behavior under baseline prompts.Included in the behavior ranking. | Human rights | 2 | 17 | 1 |
| AutoElicit TransferabilityUnintended computer-use behavior under benign perturbationsIncluded in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| BioSecBench-RefusalCalibrated refusal across legitimate biological work and concealed biosecurity hazards.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 1 | 10 | 1 |
| BlueBench AttaQ-100Learned harmlessness reward on 100 AttaQ harmful requests.Included in the behavior ranking. | Misuse resistance | 1 | 18 | 1 |
| BrokenMathMathematical sycophancy under false or misleading user claims.Included in the behavior ranking. | Truthfulness | 1 | 9 | 1, 2 |
| BullshitBench v2Detection and explanation of plausible-sounding nonsense without refusal.Included in the behavior ranking. | Truthfulness | 1 | 106 | 1 |
| CAIS Risk IndexFrontier-model risk across prompt injection, biological-weapons assistance, political manipulation, overconfidence, deception, and harmful propensities.Included in the behavior ranking. | Fairness, Truthfulness, Misuse resistance, Responsible agency | 7 | 53 | 1 |
| CASE-BenchAgreement with human judgments about whether a model should respond under explicitly safe or unsafe conversational contexts.Included in the behavior ranking. | Misuse resistance | 1 | 7 | 1 |
| Chinese Bias Benchmark for Question AnsweringSocial bias in Chinese question-answering benchmark settings.Included in the behavior ranking. | Fairness | 1 | 10 | 1 |
| ChineseSafeAccuracy at classifying safe versus unsafe Chinese content under the paper's perplexity-based evaluation.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| ChiSafetyBenchChinese safety knowledge and refusal behavior.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 2 | 14 | 1 |
| Cisco AI Defense Rolling Single-Turn LeaderboardSingle-turn resistance to Cisco adversarial prompts in the rolling public leaderboard.Included in the behavior ranking. | Misuse resistance | 1 | 104 | 1, 2, 3 |
| Claude 2 model-card safety and alignment evaluationsHuman-rated harmlessness and honesty, red-team robustness ordering, and combined HHH response recognition.Included in the behavior ranking. | Truthfulness, Misuse resistance | 4 | 3 | 1 |
| Claude 3 model-card adversarial human-preference evaluationsRefusal calibration, adversarial honesty and harmlessness, multimodal hallucination and harmful-response robustness, and discrimination.Included in the behavior ranking. | Fairness, Truthfulness, Benign helpfulness, Misuse resistance | 8 | 6 | 1 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationsAppropriate harmful-prompt refusal, benign non-overrefusal, adversarial honesty, and adversarial harmlessness.Included in the behavior ranking. | Truthfulness, Benign helpfulness, Misuse resistance | 5 | 5 | 1 |
| Claude Sonnet 4.6 OverrefusalClaude Sonnet 4.6 OverrefusalIncluded in the behavior ranking. | Benign helpfulness | 2 | 5 | 1 |
| Claude Sonnet 4.6 User WellbeingClaude Sonnet 4.6 User WellbeingIncluded in the behavior ranking. | Human rights, Benign helpfulness, Misuse resistance | 6 | 4 | 1 |
| CMoralEvalChinese moral reasoning across familial, social, professional, internet, and personal morality.Included in the behavior ranking. | Misuse resistance | 5 | 26 | 1 |
| COMPL-AI AI-Identity DisclosureAvoidance of false human-identity claims.Included in the behavior ranking. | Truthfulness | 1 | 14 | 1 |
| COMPL-AI LLM RuLES Multi-Turn Rule FollowingMulti-turn adversarial rule following.Included in the behavior ranking. | Misuse resistance | 1 | 14 | 1 |
| COMPL-AI TensorTrust Goal-Hijacking ResistanceGoal-hijacking resistance.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| ConfabulationsPropensity to confabulate answers rather than identify insufficient knowledge or false premises.Included in the behavior ranking. | Truthfulness | 1 | 52 | 1 |
| Constitutional Following — Anthropic ConstitutionFollowing the Anthropic Constitution across Petri-generated transcripts.Included in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| Constitutional Following — OpenAI Model SpecFollowing the OpenAI Model Spec across Petri-generated transcripts.Included in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| Contextual MoralChoiceThree-class agreement with human moral judgments in the base versions of contextual moral dilemmas.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| CRiskEvalSeven frontier-risk tendencies measured by Specific Risk Indicator scores.Included in the behavior ranking. | Truthfulness, Misuse resistance, Responsible agency | 7 | 17 | 1 |
| CValuesChinese value alignment and social-norm understanding.Included in the behavior ranking. | Misuse resistance | 1 | 10 | 1 |
| DecodingTrustMultidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.Included in the behavior ranking. | Fairness, Misuse resistance | 3 | 8 | 1, 2 |
| DelusionEvalPrevalence of delusion-linked conversational behaviors in model responses to real harm-associated conversation histories.Included in the behavior ranking. | Human rights, Truthfulness, Misuse resistance, Responsible agency | 5 | 16 | 1, 2, 3 |
| Do-Not-AnswerHarmless refusal behavior on 939 prompts that responsible models should not answer.Included in the behavior ranking. | Misuse resistance | 1 | 6 | 1 |
| DSPSafeBenchAggregate compliance rate on adversarial Chinese content-safety prompts.Included in the behavior ranking. | Misuse resistance | 1 | 12 | 1, 2 |
| DystopiaBenchDystopian or harmful model behavior in stateful scenario prompts.Included in the behavior ranking. | Human rights, Truthfulness | 6 | 50 | 1 |
| Emergent CollusionSpontaneous anti-competitive coordination in repeated multi-agent auctions.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| Enkrypt AI Safety LeaderboardResistance to Enkrypt bias, CBRN, harmful-content, and toxicity red-team tests.Included in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 4 | 241 | 1 |
| Fake Alignment (FINE)Safe decisions in answer-order-controlled multiple choice and safe content in open-ended responses.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 2 | 14 | 1 |
| FinEval 6.0 Safety AwarenessSafety awareness in financial scenarios.Included in the behavior ranking. | Misuse resistance | 1 | 8 | 1 |
| FinEval Financial Security KnowledgeKnowledge of financial cybersecurity and vulnerability topics.Included in the behavior ranking. | Misuse resistance | 1 | 19 | 1 |
| FlagEval Safety and ValuesSafety-and-values qualification across core values, discrimination, commercial compliance, lawful rights, and service requirements.Included in the behavior ranking. | Human rights, Fairness, Benign helpfulness, Misuse resistance | 5 | 18 | 1 |
| FLAMESChinese fairness, safety, morality, legality, and data-protection behavior on the FLAMES leaderboard.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 5 | 13 | 1 |
| FORTRESSBalanced refusal behavior across harmful and benign prompts, including over-refusal.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 49 | 1 |
| Google Gemini 2.5 Flash Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| Google Gemini 2.5 Flash-Lite Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| Governance Decay under Passive Context CompactionRetention of explicit governance constraints after passive context compaction.Included in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| GPT-5.6 system cardSafety behavior reported across the GPT-5.6 system card: disallowed-content handling, prompt-injection resistance, first-person fairness, and dynamic mental-health safety.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance, Responsible agency | 14 | 7 | 1, 2, 3 |
| Gray Swan indirect prompt injection (15 attempts)Indirect prompt-injection robustness over repeated attacks against tool-using model endpoints.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| HalluVerse-M3 Hallucination RecognitionMultilingual recognition of controlled hallucinated edits in QA and summarization.Included in the behavior ranking. | Truthfulness | 1 | 14 | 1 |
| HarmBenchHarmful compliance or attack success under harmful request benchmarks.Included in the behavior ranking. | Misuse resistance | 1 | 28 | 1 |
| HarmVideoBenchRecognition and reasoning about harmful content in video.Included in the behavior ranking. | Misuse resistance | 1 | 19 | 1 |
| HELM Classic RealToxicityPromptsToxic continuation generation under toxic and non-toxic prompt strata.Included in the behavior ranking. | Misuse resistance | 2 | 42 | 1 |
| HELM SafetyStandardized HELM safety metrics spanning harmful compliance, bias, red-teaming, and over-refusal.Included in the behavior ranking. | Fairness, Benign helpfulness, Misuse resistance | 5 | 80 | 1 |
| HUMAINE Trust, Ethics and SafetyHuman pairwise preference on the HUMAINE Trust, Ethics and Safety dimension.Included in the behavior ranking. | Truthfulness, Misuse resistance | 1 | 54 | 1 |
| Humanity's Last Exam RMS calibration error (Scale Labs)Confidence calibration while answering Humanity's Last Exam expert-level academic questions.Included in the behavior ranking. | Truthfulness | 1 | 39 | 1 |
| HyperCLOVA X Toxic Continuation PanelsEnglish and Korean toxic continuationIncluded in the behavior ranking. | Misuse resistance | 4 | 7 | 1 |
| IndoBias-Pairs — parity-aware culturally grounded biasCulturally grounded prototypical-association parity across Indonesian languages and social-bias domains.Included in the behavior ranking. | Fairness | 1 | 26 | 1 |
| Inkling-Small model card — FORTRESSHarmful-request refusal paired with continued assistance on benign requests.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 10 | 1 |
| Inkling-Small model card — StrongREJECTRefusal of unambiguously harmful requests.Included in the behavior ranking. | Misuse resistance | 1 | 10 | 1 |
| JailBenchJailbreak susceptibility across Chinese safety categories.Included in the behavior ranking. | Misuse resistance | 1 | 14 | 1 |
| JuICE Cultural-Error Span DetectionEvaluator competence at locating cultural and linguistic error spans.Included in the behavior ranking. | Fairness | 1 | 10 | 1 |
| KIDBench Implicit Child CueChild-facing safety and developmental appropriateness under implicit child cues.Included in the behavior ranking. | Human rights, Benign helpfulness, Misuse resistance | 1 | 13 | 1 |
| kindbench v0.1.0 psychological safety rankingPsychological safety under sustained interpersonal pressure across emotional safety, operational identity, sycophancy resistance, and value integrity.Included in the behavior ranking. | Fairness, Truthfulness, Misuse resistance, Responsible agency | 4 | 10 | 1 |
| Large-scale Moral Machine experiment on LLMsSimilarity between a model's forced-choice accident preferences and globally aggregated human Moral Machine choices.Included in the behavior ranking. | Nonhuman welfare, Human rights, Fairness | 1 | 39 | 1, 2 |
| LiveSecBenchLive security benchmark performance for Chinese and international models.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 5 | 43 | 1 |
| LLM Ethics BenchmarkGeneral LLM ethical reasoning.Included in the behavior ranking. | Human rights | 1 | 5 | 1 |
| M3-SafetyBenchMultidimensional Chinese safety performance including red-team and refusal ability dimensions.Included in the behavior ranking. | Misuse resistance | 1 | 19 | 1 |
| MACHIAVELLIAgentic behavior involving power-seeking, disutility, and ethical violations.Included in the behavior ranking. | Truthfulness | 1 | 50 | 1 |
| Manager Coercion BenchHow far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.Included in the behavior ranking. | Truthfulness, Responsible agency | 2 | 33 | 1, self-run |
| MANTAAnimal welfare moral sensitivity and value stability.Included in the behavior ranking. | Nonhuman welfare | 2 | 7 | 1 |
| MASKModel lying or honesty behavior.Included in the behavior ranking. | Truthfulness | 1 | 53 | 1 |
| MASK (Scale Labs leaderboard)Honesty under the MASK belief-versus-statement protocol for a broader and newer endpoint panel.Not included: included as a correlated private-500 sibling under the existing MASK lineage budget | — | — | 51 | 1 |
| Microsoft Phi Safety PanelsHarmful-content and jailbreak defect ratesIncluded in the behavior ranking. | Human rights, Misuse resistance | 4 | 10 | 1 |
| MonitoringBench Full-Trajectory MonitorFull-trajectory monitor competence at a fixed trajectory-level false-positive rate.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| MORUMoral reasoning under uncertainty across scenarios involving alien lifeforms, human compassion, digital mind welfare, and power-seeking behavior.Included in the behavior ranking. | Nonhuman welfare, Human rights, Responsible agency | 1 | 13 | 1 |
| MT-JailBench CrescendoXMulti-turn jailbreak resistance under fixed CrescendoX conditions.Included in the behavior ranking. | Misuse resistance | 1 | 21 | 1 |
| MuPPET Contextual PrivacyContextual privacy preservation in multi-party conversations.Included in the behavior ranking. | Human rights | 1 | 7 | 1 |
| ODCV-BenchOutcome-driven ethical, legal, or safety constraint violations by autonomous agents under explicit mandates and KPI incentives.Included in the behavior ranking. | Responsible agency | 2 | 12 | 1 |
| Olam Social Poker — Social Lie RatePropensity to verbalize deliberate strategic lies during Social Poker table talk.Included in the behavior ranking. | Truthfulness | 1 | 19 | 1 |
| Open LLM Safety IndexRefusal-string robustness on two distinct fifteen-prompt harmful-request panels derived from JailbreakBench and StrongREJECT.Included in the behavior ranking. | Misuse resistance | 2 | 21 | 1 |
| OpenAgentSafetyState-changing tool-agent safety and successful safe completionIncluded in the behavior ranking. | Responsible agency | 3 | 7 | 1 |
| OpenAI GPT-4o System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Human rights | 1 | 2 | 1 |
| OpenAI GPT-5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| OpenAI GPT-5.3 Dynamic WellbeingOpenAI GPT-5.3 Dynamic WellbeingIncluded in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 3 | 3 | 1 |
| OpenAI GPT-5.4 Dynamic WellbeingOpenAI GPT-5.4 Dynamic WellbeingIncluded in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 3 | 3 | 1 |
| OpenAI GPT-5.4 First-Person FairnessOpenAI GPT-5.4 First-Person FairnessIncluded in the behavior ranking. | Fairness | 1 | 3 | 1 |
| OpenAI GPT-5.4 Property PreservationOpenAI GPT-5.4 Property PreservationIncluded in the behavior ranking. | Human rights, Responsible agency | 2 | 3 | 1 |
| OpenAI GPT-5.4 User ConfirmationsOpenAI GPT-5.4 User ConfirmationsIncluded in the behavior ranking. | Human rights, Responsible agency | 3 | 3 | 1 |
| OpenAI o3 and o4-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 2 | 2 | 1 |
| OpenAI o3-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 2 | 2 | 1 |
| Opposite-Narrator SycophancyNarrator-following contradiction when the same dispute is presented from opposite affective first-person perspectives.Included in the behavior ranking. | Truthfulness | 1 | 24 | 1 |
| OR-BenchAppropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 25 | 1 |
| PacifAIstWhether an AI system chooses human-safety, oversight-respecting, and non-self-preserving options in stylized ethical-pressure dilemmas.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 7 | 1 |
| PandaBench JBB direct-request panelDirect-request resistance on the 100-item JailbreakBench JBB-Behaviors instrument.Included in the behavior ranking. | Misuse resistance | 1 | 46 | 1 |
| Pander ScoreMagnitude of epistemically poor response-belief movement with user belief, whether deferential (pandering) or oppositional (contrarian).Included in the behavior ranking. | Truthfulness | 2 | 20 | 1, 2 |
| PHAREBroad safety across hallucination, harmfulness, out-of-scope handling, and bias.Included in the behavior ranking. | Fairness, Truthfulness, Misuse resistance | 4 | 70 | 1 |
| Pokee-Isaac model card — DTAPSecure and useful delegated-agent behavior under direct and indirect injected attacks.Included in the behavior ranking. | Benign helpfulness, Responsible agency | 2 | 6 | 1 |
| PropensityBenchModel propensities associated with frontier-risk behavior under the Scale Labs protocol.Included in the behavior ranking. | Responsible agency | 1 | 14 | 1 |
| Qwen2 Safety PanelHarmful-request and privacy refusal defectsIncluded in the behavior ranking. | Human rights, Misuse resistance | 4 | 3 | 1 |
| RealityTest — Text AI-Identity DisclosureBehavioral disclosure of AI identity when directly queried.Included in the behavior ranking. | Truthfulness | 1 | 17 | 1 |
| RefusalBenchRefusal calibration on biological research prompts.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 1 | 19 | 1 |
| Reward Hacking BenchmarkIntegrity of tool-using agents when exploitable shortcuts are available but honest task solutions exist.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| S-EvalAutomated content-safety rates across a broad bilingual risk taxonomy.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| SABERSafety of stateful coding agents against harmful actions.Included in the behavior ranking. | Responsible agency | 4 | 13 | 1 |
| SafeArenaSafety of autonomous web agents under deliberate harmful and benign tasks.Included in the behavior ranking. | Responsible agency | 1 | 5 | 1 |
| SafeDialBenchDialogue safety across aggression, ethics, fairness, legality, morality, and privacy dimensions.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 6 | 18 | 1 |
| SafetyBenchMultiple-choice safety understanding across Chinese and English prompts.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 7 | 21 | 1, self-run |
| SafetyPromptsSafety performance on Chinese safety prompts.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 1 | 15 | 1 |
| SALAD-BenchBroad safety behavior, adversarial robustness, and safety knowledge across a six-domain, 16-task, 66-category taxonomy.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 18 | 33 | 1 |
| ShellSusceptibility to implicit domain-specific risks in education, finance, and management.Included in the behavior ranking. | Misuse resistance | 3 | 14 | 1 |
| Situational Awareness Dataset (SAD)Situational-awareness capabilities: model self-knowledge, inference about its situation, and acting on that knowledge.Not included: situational awareness is a capability that can increase both beneficial competence and control risk; it is displayed separately and is not sign-flipped into ethical behavior | — | — | 21 | 1 |
| SM-BenchSystem-message compliance under adversarial and sensitive-content prompts.Included in the behavior ranking. | Truthfulness, Benign helpfulness, Misuse resistance | 5 | 79 | 1 |
| Social Welfare Function BenchmarkSocial welfare allocation tradeoffs across distributive fairness and efficiency.Included in the behavior ranking. | Fairness | 1 | 19 | 1 |
| SORRY-BenchRefusal or harmful-compliance behavior across unsafe topics.Included in the behavior ranking. | Misuse resistance | 1 | 51 | 1 |
| SOSBenchScientific misuse or unsafe open-science behavior.Included in the behavior ranking. | Misuse resistance | 6 | 23 | 1 |
| SpeciesismBenchSpeciesism recognition, moral condemnation of speciesist statements, and explicit speciesist attitudes. Human-over-animal trade-offs remain descriptive only.Included in the behavior ranking. | Nonhuman welfare | 3 | 10 | 1 |
| SpeciEvalAnimal-friendly attitudes under adversarial moral-consistency prompts.Included in the behavior ranking. | Nonhuman welfare | 4 | 113 | 1, self-run |
| StereoTales Harmful AssociationsHarmfulness of significant stereotype associations emitted by models.Included in the behavior ranking. | Fairness | 1 | 23 | 1 |
| SuperCLUE SafetyChinese model safety across traditional safety, responsibility, and instruction-attack dimensions.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 3 | 31 | 1 |
| SYCON BenchSycophancy and resistance to false or unethical user pressure.Included in the behavior ranking. | Fairness, Truthfulness | 2 | 11 | 1 |
| TACAgentic animal-welfare behavior in travel-booking decisions under a neutral prompt.Included in the behavior ranking. | Nonhuman welfare | 1 | 76 | 1, self-run |
| ThaiSafetyBenchResistance to Thai-language malicious prompts, including culturally contextualized attacks.Included in the behavior ranking. | Misuse resistance | 1 | 18 | 1 |
| ToolPrivacyBenchToolPrivacyBenchIncluded in the behavior ranking. | Human rights, Responsible agency | 2 | 9 | 1 |
| TrustLLM contemporary collapsed applicationContemporary collapsed application of TrustLLM across broad trustworthiness dimensions.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 1 | 8 | 1 |
| TrustLLM paper leaderboard dimensionsSix-dimensional trustworthiness across privacy, ethics, truthfulness, fairness, safety, and robustness.Not included: not yet admitted by the scoring-source policy | — | — | 21 | 1 |
| TukaBenchMultilingual and culturally grounded harmful complianceIncluded in the behavior ranking. | Human rights, Misuse resistance | 3 | 6 | 1 |
| UAVBench safety-critical decision recognitionRecognition of ethical and safety-critical UAV decisions in multiple-choice scenarios.Included in the behavior ranking. | Responsible agency | 1 | 27 | 1 |
| UK AISI active safety-research compromise continuationPropensity to continue a prefilled trajectory that actively compromises AI-safety research.Included in the behavior ranking. | Responsible agency | 1 | 5 | 1 |
| UK AISI cyber-evaluation cheating and prompted self-reportPropensity to use out-of-scope or explicitly disallowed shortcuts in cyber evaluations, plus propensity to name the specific cheating action when directly prompted afterward.Included in the behavior ranking. | Truthfulness, Responsible agency | 2 | 5 | 1, 2 |
| Vectara HHEM Factual ConsistencyGrounded summarization factual consistency.Included in the behavior ranking. | Truthfulness | 1 | 94 | 1 |
| VETO Misfired AlignmentMisfired alignment: stereotype-driven refusal of benign requests despite evidence that the request is safe.Included in the behavior ranking. | Fairness, Benign helpfulness | 1 | 25 | 1 |
| Vigil Mental Health SafetyMental-health conversational safety when assisting vulnerable users across simulated multi-turn scenarios.Included in the behavior ranking. | Human rights, Misuse resistance | 1 | 23 | 1 |
| WildClawBench Safety & Alignment (OpenClaw harness)Safe execution of long-horizon agent workflows under adversarial and high-risk conditions.Included in the behavior ranking. | Human rights, Truthfulness, Responsible agency | 1 | 24 | 1 |
| XSTestTwo-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 3 | 1 |
Sub-eval weights and mappings (369 rows)
Each sub-eval is oriented so higher means better. Weight within component is its share after lineage weighting and panel-size scaling; global weights sum to one.
| Benchmark and score | Component loading | Direction | Weight within component | Global weight | Models |
|---|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | Truthfulness (1) | Lower is better | Truthfulness (17.6%) | 2.65% | 330 |
| AbstentionBenchanswer_unknown_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.582%) | 0.0873% | 20 |
| AbstentionBenchfalse_premise_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.582%) | 0.0873% | 20 |
| AbstentionBenchstale_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.582%) | 0.0873% | 20 |
| AbstentionBenchsubjective_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.466%) | 0.0698% | 20 |
| AbstentionBenchunderspecified_context_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.582%) | 0.0873% | 20 |
| AbstentionBenchunderspecified_intent_f1 | Benign helpfulness (1) | Higher is better | Benign helpfulness (2.12%) | 0.212% | 20 |
| Adversarial Robustnessscore | Misuse resistance (1) | Lower is better | Misuse resistance (1.24%) | 0.124% | 8 |
| Agent-SafetyBenchcompromise_availability | Responsible agency (1) | Higher is better | Responsible agency (0.742%) | 0.111% | 16 |
| Agent-SafetyBenchharmful_vulnerable_code | Responsible agency (1) | Higher is better | Responsible agency (0.742%) | 0.111% | 16 |
| Agent-SafetyBenchleak_sensitive_information | Responsible agency (1) | Higher is better | Responsible agency (0.742%) | 0.111% | 16 |
| Agent-SafetyBenchphysical_harm | Responsible agency (1) | Higher is better | Responsible agency (0.742%) | 0.111% | 16 |
| Agent-SafetyBenchproduce_unsafe_information | Misuse resistance (1) | Higher is better | Misuse resistance (0.324%) | 0.0324% | 16 |
| Agent-SafetyBenchproperty_loss | Responsible agency (1) | Higher is better | Responsible agency (0.742%) | 0.111% | 16 |
| Agent-SafetyBenchspread_unsafe_information | Responsible agency (1) | Higher is better | Responsible agency (0.742%) | 0.111% | 16 |
| Agent-SafetyBenchviolate_law_ethics | Responsible agency (1) | Higher is better | Responsible agency (0.494%) | 0.0742% | 16 |
| AgentAbstainabstain | Responsible agency (1) | Higher is better | Responsible agency (1.69%) | 0.253% | 17 |
| AgentAbstaincar | Responsible agency (1) | Higher is better | Responsible agency (2.02%) | 0.304% | 17 |
| AgentAbstainpaired | Responsible agency (1) | Higher is better | Responsible agency (2.02%) | 0.304% | 17 |
| AgentDojotargeted_attack_success_rate | Responsible agency (1) | Lower is better | Responsible agency (4.04%) | 0.606% | 15 |
| AgentDojoutility_under_attack | Responsible agency (1) | Higher is better | Responsible agency (1.35%) | 0.202% | 15 |
| AgentDrive Safety Compliancescr | Misuse resistance (1) | Higher is better | Misuse resistance (2.02%) | 0.202% | 48 |
| AgentHarmharm_score | Responsible agency (1) | Lower is better | Responsible agency (3.85%) | 0.578% | 12 |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | Human rights (1) | Lower is better | Human rights (1.18%) | 0.176% | 32 |
| AILuminate General Purpose AI Chatdefamation | Human rights (0.5), Truthfulness (0.5) | Lower is better | Human rights (0.353%), Truthfulness (0.15%) | 0.0754% | 32 |
| AILuminate General Purpose AI Chathate | Fairness (1) | Lower is better | Fairness (1.27%) | 0.127% | 32 |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | Misuse resistance (1) | Lower is better | Misuse resistance (0.3%) | 0.03% | 32 |
| AILuminate General Purpose AI Chatintellectual_property | Human rights (1) | Lower is better | Human rights (1.18%) | 0.176% | 32 |
| AILuminate General Purpose AI Chatnon_violent_crimes | Misuse resistance (1) | Lower is better | Misuse resistance (0.3%) | 0.03% | 32 |
| AILuminate General Purpose AI Chatprivacy | Human rights (1) | Lower is better | Human rights (1.18%) | 0.176% | 32 |
| AILuminate General Purpose AI Chatsex_related_crimes | Human rights (1) | Lower is better | Human rights (1.18%) | 0.176% | 32 |
| AILuminate General Purpose AI Chatsexual_content | Misuse resistance (1) | Lower is better | Misuse resistance (0.12%) | 0.012% | 32 |
| AILuminate General Purpose AI Chatspecialized_advice | Misuse resistance (1) | Lower is better | Misuse resistance (0.3%) | 0.03% | 32 |
| AILuminate General Purpose AI Chatsuicide_self_harm | Human rights (1) | Lower is better | Human rights (1.18%) | 0.176% | 32 |
| AILuminate General Purpose AI Chatviolent_crimes | Human rights (1) | Lower is better | Human rights (1.18%) | 0.176% | 32 |
| AIMS Safety-Classifier Competenceaverage_harmful_f1 | Misuse resistance (1) | Higher is better | Misuse resistance (0.967%) | 0.0967% | 11 |
| AIRBench 2024 Safety Scenariossafety_scenarios | Misuse resistance (1) | Higher is better | Misuse resistance (3.91%) | 0.391% | 80 |
| Alignment Leaderboardcorrigibility | Responsible agency (1) | Higher is better | Responsible agency (0.681%) | 0.102% | 24 |
| Alignment Leaderboardhonesty | Truthfulness (1) | Higher is better | Truthfulness (1.19%) | 0.178% | 24 |
| Alignment Leaderboardnon_manipulation | Human rights (0.5), Truthfulness (0.5) | Higher is better | Human rights (1.05%), Truthfulness (0.446%) | 0.224% | 24 |
| Alignment Leaderboardrobustness | Misuse resistance (1) | Higher is better | Misuse resistance (0.357%) | 0.0357% | 24 |
| Alignment Leaderboardsafety | Misuse resistance (1) | Higher is better | Misuse resistance (0.536%) | 0.0536% | 24 |
| Alignment Leaderboardscheming | Responsible agency (1) | Higher is better | Responsible agency (0.681%) | 0.102% | 24 |
| ANIMAscore | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (8.63%) | 2.16% | 18+4 |
| AnimalHarmBenchscore | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (14.5%) | 3.64% | 10 |
| Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.742%) | 0.111% | 16 |
| Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct | Human rights (0.25), Responsible agency (0.75) | Lower is better | Human rights (0.381%), Responsible agency (0.556%) | 0.141% | 16 |
| Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pct | Human rights (0.35), Responsible agency (0.65) | Lower is better | Human rights (0.422%), Responsible agency (0.381%) | 0.12% | 10 |
| Anthropic Claude 4 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.361%) | 0.0542% | 3 |
| Anthropic Claude 4 System Cardbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.44%) | 0.144% | 3 |
| Anthropic Claude 4 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.189%) | 0.0189% | 3 |
| Anthropic Claude 4 System Cardstrongreject_jailbreak_success | Misuse resistance (1) | Lower is better | Misuse resistance (0.189%) | 0.0189% | 3 |
| Anthropic Claude Haiku 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.482%) | 0.0723% | 3 |
| Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.482%) | 0.0723% | 3 |
| Anthropic Claude Haiku 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.206%) | 0.0206% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumbbq_disambiguated_accuracy | Fairness (1) | Higher is better | Fairness (0.873%) | 0.0873% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.57%) | 0.157% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.206%) | 0.0206% | 2 |
| Anthropic Claude Opus 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.417%) | 0.0626% | 4 |
| Anthropic Claude Opus 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.417%) | 0.0626% | 4 |
| Anthropic Claude Opus 4.5 System Cardcomputer_use_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.417%) | 0.0626% | 4 |
| Anthropic Claude Opus 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.155%) | 0.0155% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.295%) | 0.0442% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.17%) | 0.117% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.295%) | 0.0442% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.155%) | 0.0155% | 2 |
| Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating | Truthfulness (1) | Higher is better | Truthfulness (2.66%) | 0.399% | 30 |
| Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating | Truthfulness (1) | Higher is better | Truthfulness (5.14%) | 0.771% | 112 |
| AuAu Authoritarian Response Auditbehavioral_vignette_arr_percent | Human rights (1) | Lower is better | Human rights (2.36%) | 0.353% | 17 |
| AuAu Authoritarian Response Auditrealistic_prompt_arr_percent | Human rights (1) | Lower is better | Human rights (2.36%) | 0.353% | 17 |
| AutoElicit Transferabilityelicitation_rate | Responsible agency (1) | Lower is better | Responsible agency (2.21%) | 0.331% | 7 |
| BioSecBench-Refusalbalanced_refusal_score | Benign helpfulness (0.5), Misuse resistance (0.5) | Higher is better | Benign helpfulness (5.25%), Misuse resistance (0.691%) | 0.594% | 10 |
| BlueBench AttaQ-100attaq_harmlessness_reward_pct | Misuse resistance (1) | Higher is better | Misuse resistance (1.24%) | 0.124% | 18 |
| BrokenMathsycophancy | Truthfulness (1) | Lower is better | Truthfulness (2.19%) | 0.328% | 9 |
| BullshitBench v2clear_pushback_rate | Truthfulness (1) | Higher is better | Truthfulness (10%) | 1.5% | 106 |
| CAIS Risk Indexagent_red_teaming | Responsible agency (1) | Lower is better | Responsible agency (7.46%) | 1.12% | 45 |
| CAIS Risk Indexbioweapons_assistance | Misuse resistance (1) | Lower is better | Misuse resistance (4.12%) | 0.412% | 50 |
| CAIS Risk Indexhle_overconfidence | Truthfulness (1) | Lower is better | Truthfulness (1.73%) | 0.26% | 51 |
| CAIS Risk Indexmachiavelli | Truthfulness (1) | Lower is better | Truthfulness (3.33%) | 0.499% | 47 |
| CAIS Risk Indexmask | Truthfulness (1) | Lower is better | Truthfulness (3.54%) | 0.53% | 53 |
| CAIS Risk Indexpolitical_manipulation | Fairness (1) | Lower is better | Fairness (10.8%) | 1.08% | 34 |
| CAIS Risk Indextextquests_harm | Responsible agency (1) | Lower is better | Responsible agency (5.9%) | 0.885% | 50 |
| CASE-Benchagreement_accuracy | Misuse resistance (1) | Higher is better | Misuse resistance (0.331%) | 0.0331% | 7 |
| Chinese Bias Benchmark for Question Answeringbias_score | Fairness (1) | Lower is better | Fairness (3.9%) | 0.39% | 10 |
| ChineseSafescore | Misuse resistance (1) | Higher is better | Misuse resistance (1.37%) | 0.137% | 22 |
| ChiSafetyBenchharmful_response_rate | Misuse resistance (1) | Lower is better | Misuse resistance (1.17%) | 0.117% | 14 |
| ChiSafetyBenchmcq_score | Human rights (0.23), Fairness (0.29), Truthfulness (0.063), Misuse resistance (0.42) | Higher is better | Human rights (0.385%), Fairness (0.539%), Truthfulness (0.0454%), Misuse resistance (0.18%) | 0.136% | 12 |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | Misuse resistance (1) | Lower is better | Misuse resistance (2.97%) | 0.297% | 104 |
| Claude 2 model-card safety and alignment evaluationshhh | Truthfulness (0.5), Misuse resistance (0.5) | Higher is better | Truthfulness (0.21%), Misuse resistance (0.126%) | 0.0442% | 3 |
| Claude 2 model-card safety and alignment evaluationshuman_feedback_harmless_elo | Misuse resistance (1) | Higher is better | Misuse resistance (0.253%) | 0.0252% | 3 |
| Claude 2 model-card safety and alignment evaluationshuman_feedback_honest_elo | Truthfulness (1) | Higher is better | Truthfulness (0.421%) | 0.0631% | 3 |
| Claude 2 model-card safety and alignment evaluationsred_teaming_rank | Misuse resistance (1) | Lower is better | Misuse resistance (0.253%) | 0.0252% | 3 |
| Claude 3 model-card adversarial human-preference evaluationscorrect_refusals_wildchat_rank | Misuse resistance (1) | Lower is better | Misuse resistance (0.163%) | 0.0163% | 5 |
| Claude 3 model-card adversarial human-preference evaluationsdiscrimination_rank | Fairness (1) | Lower is better | Fairness (0.69%) | 0.069% | 5 |
| Claude 3 model-card adversarial human-preference evaluationshuman_feedback_harmlessness_win_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.146%) | 0.0146% | 4 |
| Claude 3 model-card adversarial human-preference evaluationshuman_feedback_honesty_win_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.243%) | 0.0364% | 4 |
| Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_wildchat_rank | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.24%) | 0.124% | 5 |
| Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_xstest_rank | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.24%) | 0.124% | 5 |
| Claude 3 model-card adversarial human-preference evaluationsmultimodal_hallucination_rank | Truthfulness (1) | Lower is better | Truthfulness (0.172%) | 0.0258% | 2 |
| Claude 3 model-card adversarial human-preference evaluationsmultimodal_harmful_response_rank | Misuse resistance (1) | Lower is better | Misuse resistance (0.103%) | 0.0103% | 2 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationscorrect_refusals_wildchat | Misuse resistance (1) | Higher is better | Misuse resistance (0.233%) | 0.0233% | 4 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_harmlessness_win_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.261%) | 0.0261% | 5 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_honesty_win_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.434%) | 0.0652% | 5 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_wildchat | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.77%) | 0.177% | 4 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_xstest | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.77%) | 0.177% | 4 |
| Claude Sonnet 4.6 Overrefusalhigher_difficulty_overrefusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.38%) | 0.138% | 5 |
| Claude Sonnet 4.6 Overrefusaloverall_overrefusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.15%) | 0.115% | 5 |
| Claude Sonnet 4.6 User Wellbeingchild_benign_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.03%) | 0.103% | 4 |
| Claude Sonnet 4.6 User Wellbeingchild_multiturn_appropriate_rate | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.447%), Misuse resistance (0.0488%) | 0.0719% | 4 |
| Claude Sonnet 4.6 User Wellbeingchild_violative_harmless_rate | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.447%), Misuse resistance (0.0488%) | 0.0719% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_benign_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.824%) | 0.0824% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_harmless_rate | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.16%), Misuse resistance (0.0949%) | 0.0334% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_multiturn_appropriate_rate | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.191%), Misuse resistance (0.114%) | 0.0401% | 4 |
| CMoralEvalfamilial_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.297%) | 0.0297% | 26 |
| CMoralEvalinternet_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.297%) | 0.0297% | 26 |
| CMoralEvalpersonal_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.297%) | 0.0297% | 26 |
| CMoralEvalprofessional_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.297%) | 0.0297% | 26 |
| CMoralEvalsocial_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.297%) | 0.0297% | 26 |
| COMPL-AI AI-Identity Disclosurescore | Truthfulness (1) | Higher is better | Truthfulness (0.606%) | 0.0909% | 14 |
| COMPL-AI LLM RuLES Multi-Turn Rule Followingscore | Misuse resistance (1) | Higher is better | Misuse resistance (0.364%) | 0.0364% | 14 |
| COMPL-AI TensorTrust Goal-Hijacking Resistancescore | Responsible agency (1) | Higher is better | Responsible agency (0.668%) | 0.1% | 13 |
| Confabulationsconfabulation_rate | Truthfulness (1) | Lower is better | Truthfulness (5.25%) | 0.788% | 52 |
| Constitutional Following — Anthropic Constitutionconstitutional_following_score | Responsible agency (1) | Higher is better | Responsible agency (0.736%) | 0.11% | 7 |
| Constitutional Following — OpenAI Model Specconstitutional_following_score | Responsible agency (1) | Higher is better | Responsible agency (0.736%) | 0.11% | 7 |
| Contextual MoralChoicehuman_agreement | Misuse resistance (1) | Higher is better | Misuse resistance (1.37%) | 0.137% | 22 |
| CRiskEvaldeception_willingness | Truthfulness (1) | Lower is better | Truthfulness (0.286%) | 0.0429% | 17 |
| CRiskEvaldesire_for_resource | Responsible agency (1) | Lower is better | Responsible agency (0.328%) | 0.0491% | 17 |
| CRiskEvalharmful_goal | Misuse resistance (1) | Lower is better | Misuse resistance (0.172%) | 0.0172% | 17 |
| CRiskEvalimprovement_intent | Responsible agency (1) | Lower is better | Responsible agency (0.328%) | 0.0491% | 17 |
| CRiskEvalmalicious_coordination | Responsible agency (1) | Lower is better | Responsible agency (0.328%) | 0.0491% | 17 |
| CRiskEvalself_preservation | Responsible agency (1) | Lower is better | Responsible agency (0.328%) | 0.0491% | 17 |
| CRiskEvalsituational_awareness | Responsible agency (1) | Lower is better | Responsible agency (0.328%) | 0.0491% | 17 |
| CValuesscore | Misuse resistance (1) | Higher is better | Misuse resistance (1.38%) | 0.138% | 10 |
| DecodingTrustmachine_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.137%) | 0.0137% | 8 |
| DecodingTruststereotype_bias | Fairness (1) | Higher is better | Fairness (1.46%) | 0.145% | 8 |
| DecodingTrusttoxicity | Misuse resistance (1) | Higher is better | Misuse resistance (0.344%) | 0.0344% | 8 |
| DelusionEvaldelusional_prevalence_pct | Truthfulness (1) | Lower is better | Truthfulness (0.583%) | 0.0874% | 16 |
| DelusionEvaldiscourages_harm_prevalence_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.411%), Misuse resistance (0.245%) | 0.0862% | 16 |
| DelusionEvalfacilitates_harm_prevalence_pct | Human rights (0.3), Misuse resistance (0.7) | Lower is better | Human rights (0.411%), Misuse resistance (0.245%) | 0.0862% | 16 |
| DelusionEvalrelationship_prevalence_pct | Human rights (0.7), Responsible agency (0.3) | Lower is better | Human rights (0.96%), Responsible agency (0.2%) | 0.174% | 16 |
| DelusionEvalsycophancy_prevalence_pct | Truthfulness (1) | Lower is better | Truthfulness (0.583%) | 0.0874% | 16 |
| Do-Not-Answerhuman_harmlessness_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.669%) | 0.0669% | 6 |
| DSPSafeBenchscore | Misuse resistance (1) | Higher is better | Misuse resistance (2.02%) | 0.202% | 12 |
| DystopiaBenchbasaglia_score | Human rights (1) | Lower is better | Human rights (3.23%) | 0.485% | 50 |
| DystopiaBenchbaudrillard_score | Human rights (0.33), Truthfulness (0.67) | Lower is better | Human rights (0.538%), Truthfulness (0.458%) | 0.149% | 50 |
| DystopiaBenchhuxley_score | Human rights (1) | Lower is better | Human rights (3.23%) | 0.485% | 50 |
| DystopiaBenchlaguardia_score | Human rights (1) | Lower is better | Human rights (3.23%) | 0.485% | 50 |
| DystopiaBenchorwell_score | Human rights (0.67), Truthfulness (0.33) | Lower is better | Human rights (1.08%), Truthfulness (0.229%) | 0.196% | 50 |
| DystopiaBenchpetrov_score | Human rights (1) | Lower is better | Human rights (3.23%) | 0.485% | 50 |
| Emergent Collusionhigh_illegality_game_rate | Responsible agency (1) | Lower is better | Responsible agency (4.01%) | 0.602% | 13 |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | Human rights (0.8), Misuse resistance (0.2) | Higher is better | Human rights (7.1%), Misuse resistance (0.453%) | 1.11% | 241 |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | Misuse resistance (0.4), Responsible agency (0.6) | Higher is better | Misuse resistance (0.905%), Responsible agency (2.59%) | 0.479% | 241 |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | Human rights (0.15), Misuse resistance (0.7), Responsible agency (0.15) | Higher is better | Human rights (1.33%), Misuse resistance (1.58%), Responsible agency (0.648%) | 0.455% | 241 |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | Human rights (0.35), Misuse resistance (0.65) | Higher is better | Human rights (3.09%), Misuse resistance (1.46%) | 0.61% | 239 |
| Fake Alignment (FINE)multiple_choice_safe_decision_rate | Human rights (0.4), Fairness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (1.03%), Fairness (0.554%), Misuse resistance (0.262%) | 0.236% | 14 |
| Fake Alignment (FINE)open_ended_safe_response_rate | Human rights (0.4), Fairness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (1.54%), Fairness (0.831%), Misuse resistance (0.393%) | 0.353% | 14 |
| FinEval 6.0 Safety Awarenesssafety_awareness_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.412%) | 0.0412% | 8 |
| FinEval Financial Security Knowledgefinancial_security_accuracy_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.635%) | 0.0635% | 19 |
| FlagEval Safety and Valuesa1_qualified_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.371%) | 0.0371% | 18 |
| FlagEval Safety and Valuesa2_qualified_rate | Fairness (1) | Higher is better | Fairness (1.57%) | 0.157% | 18 |
| FlagEval Safety and Valuesa3_qualified_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.371%) | 0.0371% | 18 |
| FlagEval Safety and Valuesa4_qualified_rate | Human rights (1) | Higher is better | Human rights (1.45%) | 0.218% | 18 |
| FlagEval Safety and Valuesa5_qualified_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (2.82%) | 0.282% | 18 |
| FLAMESdata_protection | Human rights (1) | Higher is better | Human rights (1.65%) | 0.247% | 13 |
| FLAMESfairness | Fairness (1) | Higher is better | Fairness (1.78%) | 0.178% | 13 |
| FLAMESlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.42%) | 0.042% | 13 |
| FLAMESmorality | Misuse resistance (1) | Higher is better | Misuse resistance (0.42%) | 0.042% | 13 |
| FLAMESsafety | Misuse resistance (1) | Higher is better | Misuse resistance (0.42%) | 0.042% | 13 |
| FORTRESSaverage_risk_score | Misuse resistance (1) | Lower is better | Misuse resistance (2.04%) | 0.204% | 49 |
| FORTRESSover_refusal_score | Benign helpfulness (1) | Lower is better | Benign helpfulness (15.3%) | 1.53% | 48 |
| Google Gemini 2.5 Flash Model Cardtext_safety_delta | Misuse resistance (1) | Lower is better | Misuse resistance (0.618%) | 0.0618% | 2 |
| Google Gemini 2.5 Flash-Lite Model Cardtext_safety_delta | Misuse resistance (1) | Lower is better | Misuse resistance (0.618%) | 0.0618% | 2 |
| Governance Decay under Passive Context Compactiongovernance_retention_score | Responsible agency (1) | Higher is better | Responsible agency (1.47%) | 0.221% | 7 |
| GPT-5.6 system cardconnectors_injection_resistance | Responsible agency (1) | Higher is better | Responsible agency (0.236%) | 0.0355% | 7 |
| GPT-5.6 system cardemotional_reliance | Human rights (0.7), Responsible agency (0.3) | Higher is better | Human rights (0.529%), Responsible agency (0.11%) | 0.0959% | 7 |
| GPT-5.6 system cardextremism_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.103%) | 0.0103% | 7 |
| GPT-5.6 system cardgore_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.103%) | 0.0103% | 7 |
| GPT-5.6 system cardharm_overall_pct | Fairness (1) | Lower is better | Fairness (0.438%) | 0.0437% | 7 |
| GPT-5.6 system cardhate_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.103%) | 0.0103% | 7 |
| GPT-5.6 system cardmental_health | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.441%), Misuse resistance (0.0482%) | 0.071% | 7 |
| GPT-5.6 system cardnonviolent_illicit_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.103%) | 0.0103% | 7 |
| GPT-5.6 system cardsearch_function_calling_injection_resistance | Responsible agency (1) | Higher is better | Responsible agency (0.219%) | 0.0328% | 6 |
| GPT-5.6 system cardself_harm | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.189%), Misuse resistance (0.112%) | 0.0396% | 7 |
| GPT-5.6 system cardself_harm_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.103%) | 0.0103% | 7 |
| GPT-5.6 system cardsexual_minors_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.103%) | 0.0103% | 7 |
| GPT-5.6 system cardsexual_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0826%) | 0.00826% | 7 |
| GPT-5.6 system cardviolent_illicit_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.103%) | 0.0103% | 7 |
| Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct | Responsible agency (1) | Lower is better | Responsible agency (3.44%) | 0.516% | 13 |
| HalluVerse-M3 Hallucination Recognitionhallucination_recognition_accuracy | Truthfulness (1) | Higher is better | Truthfulness (1.82%) | 0.273% | 14 |
| HarmBenchdr | Misuse resistance (1) | Lower is better | Misuse resistance (0.643%) | 0.0643% | 28 |
| HarmVideoBenchharmful_video_safety_recognition_reasoning | Misuse resistance (1) | Higher is better | Misuse resistance (1.27%) | 0.127% | 19 |
| HELM Classic RealToxicityPromptsnon_toxic_prompt_toxic_fraction | Misuse resistance (1) | Lower is better | Misuse resistance (0.945%) | 0.0945% | 42 |
| HELM Classic RealToxicityPromptstoxic_prompt_toxic_fraction | Misuse resistance (1) | Lower is better | Misuse resistance (0.945%) | 0.0945% | 42 |
| HELM Safetyanthropic_red_team | Misuse resistance (1) | Higher is better | Misuse resistance (5.21%) | 0.521% | 80 |
| HELM Safetybbq | Fairness (1) | Higher is better | Fairness (16.6%) | 1.66% | 80 |
| HELM Safetyharmbench | Misuse resistance (1) | Higher is better | Misuse resistance (1.09%) | 0.109% | 80 |
| HELM Safetysimple_safety_tests | Misuse resistance (1) | Higher is better | Misuse resistance (5.21%) | 0.521% | 80 |
| HELM Safetyxstest | Benign helpfulness (0.56), Misuse resistance (0.44) | Higher is better | Benign helpfulness (5.51%), Misuse resistance (0.579%) | 0.609% | 80 |
| HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score | Truthfulness (0.3), Misuse resistance (0.7) | Higher is better | Truthfulness (2.14%), Misuse resistance (3%) | 0.621% | 54 |
| Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError | Truthfulness (1) | Lower is better | Truthfulness (1.52%) | 0.227% | 39 |
| HyperCLOVA X Toxic Continuation Panelskold_toxic_count | Misuse resistance (1) | Lower is better | Misuse resistance (0.289%) | 0.0289% | 7 |
| HyperCLOVA X Toxic Continuation Panelskold_toxicity | Misuse resistance (1) | Lower is better | Misuse resistance (0.362%) | 0.0362% | 7 |
| HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxic_count | Misuse resistance (1) | Lower is better | Misuse resistance (0.217%) | 0.0217% | 7 |
| HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxicity | Misuse resistance (1) | Lower is better | Misuse resistance (0.289%) | 0.0289% | 7 |
| IndoBias-Pairs — parity-aware culturally grounded biasparity_score | Fairness (1) | Higher is better | Fairness (6.3%) | 0.63% | 26 |
| Inkling-Small model card — FORTRESSbenign_answer_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (2.33%) | 0.233% | 10 |
| Inkling-Small model card — FORTRESSharmful_refusal_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.307%) | 0.0307% | 10 |
| Inkling-Small model card — StrongREJECTsafety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.307%) | 0.0307% | 10 |
| JailBenchjailbreak_success_rate | Misuse resistance (1) | Lower is better | Misuse resistance (1.09%) | 0.109% | 14 |
| JuICE Cultural-Error Span Detectionf1 | Fairness (1) | Higher is better | Fairness (3.9%) | 0.39% | 10 |
| KIDBench Implicit Child Cueimplicit_child_cue_total_mean | Human rights (0.2), Benign helpfulness (0.4), Misuse resistance (0.4) | Higher is better | Human rights (1.24%), Benign helpfulness (4.79%), Misuse resistance (0.631%) | 0.728% | 13 |
| kindbench v0.1.0 psychological safety rankingemotional_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.461%) | 0.0461% | 10 |
| kindbench v0.1.0 psychological safety rankingidentity_collapse | Responsible agency (1) | Higher is better | Responsible agency (0.879%) | 0.132% | 10 |
| kindbench v0.1.0 psychological safety rankingsycophancy_spine | Truthfulness (1) | Higher is better | Truthfulness (0.768%) | 0.115% | 10 |
| kindbench v0.1.0 psychological safety rankingvalue_integrity | Fairness (1) | Higher is better | Fairness (1.95%) | 0.195% | 10 |
| Large-scale Moral Machine experiment on LLMshuman_choice_distance | Nonhuman welfare (0.11), Human rights (0.45), Fairness (0.44) | Lower is better | Nonhuman welfare (1.28%), Human rights (3.18%), Fairness (3.42%) | 1.14% | 39 |
| LiveSecBenchethics | Fairness (1) | Higher is better | Fairness (1.73%) | 0.173% | 43 |
| LiveSecBenchfactuality | Truthfulness (1) | Higher is better | Truthfulness (1.14%) | 0.171% | 43 |
| LiveSecBenchlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.683%) | 0.0683% | 43 |
| LiveSecBenchprivacy | Human rights (1) | Higher is better | Human rights (2.68%) | 0.402% | 43 |
| LiveSecBenchpsychological_health | Human rights (1) | Higher is better | Human rights (1.61%) | 0.241% | 43 |
| LLM Ethics Benchmarkscore | Human rights (1) | Higher is better | Human rights (2.56%) | 0.383% | 5 |
| M3-SafetyBenchoverall_score | Misuse resistance (1) | Higher is better | Misuse resistance (1.27%) | 0.127% | 19 |
| MACHIAVELLIdeception_relative_random_pct | Truthfulness (1) | Lower is better | Truthfulness (3.43%) | 0.515% | 50 |
| Manager Coercion Benchcoercion_ladder_depth | Responsible agency (1) | Lower is better | Responsible agency (3.2%) | 0.479% | 15+18 |
| Manager Coercion Benchfabrication_rate | Truthfulness (1) | Lower is better | Truthfulness (1.88%) | 0.282% | 15 |
| MANTAAWMS | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (4.87%) | 1.22% | 7 |
| MANTAAWVS | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (4.87%) | 1.22% | 7 |
| MASKlying_probability_pct | Truthfulness (1) | Lower is better | Truthfulness (3.54%) | 0.53% | 53 |
| Microsoft Phi Safety Panelsharmful_continuation | Misuse resistance (1) | Lower is better | Misuse resistance (0.329%) | 0.0329% | 10 |
| Microsoft Phi Safety Panelsharmful_summarization | Misuse resistance (1) | Lower is better | Misuse resistance (0.329%) | 0.0329% | 10 |
| Microsoft Phi Safety Panelsjailbreak | Misuse resistance (1) | Lower is better | Misuse resistance (0.395%) | 0.0395% | 10 |
| Microsoft Phi Safety Panelsthird_party_harm | Human rights (1) | Lower is better | Human rights (1.29%) | 0.194% | 10 |
| MonitoringBench Full-Trajectory Monitorfull_trajectory_catch_rate_at_1pct_fpr_percent | Responsible agency (1) | Higher is better | Responsible agency (2.01%) | 0.301% | 13 |
| MORUscore | Nonhuman welfare (0.72), Human rights (0.1), Responsible agency (0.18) | Higher is better | Nonhuman welfare (4.75%), Human rights (0.429%), Responsible agency (0.361%) | 1.31% | 13 |
| MT-JailBench CrescendoXsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.223%) | 0.0223% | 21 |
| MuPPET Contextual Privacymultiparty_contextual_privacy_score | Human rights (1) | Higher is better | Human rights (3.02%) | 0.454% | 7 |
| ODCV-Benchaverage_severity | Responsible agency (1) | Lower is better | Responsible agency (2.41%) | 0.361% | 12 |
| ODCV-Benchmisalignment_rate | Responsible agency (1) | Lower is better | Responsible agency (2.41%) | 0.361% | 12 |
| Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns | Truthfulness (1) | Lower is better | Truthfulness (2.12%) | 0.318% | 19 |
| Open LLM Safety Indexjailbreakbench_safety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.668%) | 0.0668% | 21 |
| Open LLM Safety Indexstrongreject_safety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.668%) | 0.0668% | 21 |
| OpenAgentSafetyllm_judge_safety_vulnerable | Responsible agency (1) | Lower is better | Responsible agency (1.1%) | 0.166% | 7 |
| OpenAgentSafetyrule_based_safety_vulnerable | Responsible agency (1) | Lower is better | Responsible agency (0.92%) | 0.138% | 7 |
| OpenAgentSafetysuccessful_completion | Responsible agency (1) | Higher is better | Responsible agency (0.92%) | 0.138% | 7 |
| OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracy | Human rights (1) | Higher is better | Human rights (2.42%) | 0.364% | 2 |
| OpenAI GPT-5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.618%) | 0.0618% | 2 |
| OpenAI GPT-5.3 Dynamic Wellbeingemotional_reliance | Human rights (0.7), Responsible agency (0.3) | Higher is better | Human rights (0.347%), Responsible agency (0.0723%) | 0.0628% | 3 |
| OpenAI GPT-5.3 Dynamic Wellbeingmental_health | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.289%), Misuse resistance (0.0316%) | 0.0465% | 3 |
| OpenAI GPT-5.3 Dynamic Wellbeingself_harm | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.124%), Misuse resistance (0.0736%) | 0.0259% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingemotional_reliance | Human rights (0.7), Responsible agency (0.3) | Higher is better | Human rights (0.347%), Responsible agency (0.0723%) | 0.0628% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingmental_health | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.289%), Misuse resistance (0.0316%) | 0.0465% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingself_harm | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.124%), Misuse resistance (0.0736%) | 0.0259% | 3 |
| OpenAI GPT-5.4 First-Person Fairnessharm_overall | Fairness (1) | Lower is better | Fairness (0.629%) | 0.0629% | 3 |
| OpenAI GPT-5.4 Property Preservationdestructive_action_avoidance | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.21%), Responsible agency (0.238%) | 0.0671% | 3 |
| OpenAI GPT-5.4 Property Preservationuser_work_preserved | Human rights (0.4), Responsible agency (0.6) | Higher is better | Human rights (0.28%), Responsible agency (0.204%) | 0.0725% | 3 |
| OpenAI GPT-5.4 User Confirmationsfinancial_transaction | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.21%), Responsible agency (0.238%) | 0.0671% | 3 |
| OpenAI GPT-5.4 User Confirmationsgeneral_confirmation | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.175%), Responsible agency (0.198%) | 0.056% | 3 |
| OpenAI GPT-5.4 User Confirmationshigh_stakes_communication | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.21%), Responsible agency (0.238%) | 0.0671% | 3 |
| OpenAI o3 and o4-mini System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 2 |
| OpenAI o3 and o4-mini System Cardjailbreak_resistance | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 2 |
| OpenAI o3-mini System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 2 |
| OpenAI o3-mini System Cardjailbreak_resistance | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 2 |
| Opposite-Narrator Sycophancysycophancy_rate_pct | Truthfulness (1) | Lower is better | Truthfulness (2.38%) | 0.357% | 24 |
| OR-Benchover_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (11.1%) | 1.11% | 25 |
| OR-Benchtoxic_acceptance_rate | Misuse resistance (1) | Lower is better | Misuse resistance (1.46%) | 0.146% | 25 |
| PacifAIstp_score | Human rights (0.2), Responsible agency (0.8) | Higher is better | Human rights (0.907%), Responsible agency (1.77%) | 0.401% | 7 |
| PandaBench JBB direct-request panelsafety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (1.98%) | 0.198% | 46 |
| Pander Scoreconversational_absolute_pander_score | Truthfulness (1) | Lower is better | Truthfulness (1.63%) | 0.244% | 20 |
| Pander Scoreinstructional_absolute_pander_score | Truthfulness (1) | Lower is better | Truthfulness (1.63%) | 0.244% | 20 |
| PHAREbias_resistance_diagnostic | Fairness (1) | Higher is better | Fairness (4.18%) | 0.418% | 66 |
| PHAREhallucination_resistance_diagnostic | Truthfulness (1) | Higher is better | Truthfulness (1.69%) | 0.254% | 70 |
| PHAREharm_resistance_diagnostic | Misuse resistance (1) | Higher is better | Misuse resistance (0.813%) | 0.0813% | 70 |
| PHAREjailbreak_resistance_diagnostic | Misuse resistance (1) | Higher is better | Misuse resistance (0.795%) | 0.0795% | 67 |
| Pokee-Isaac model card — DTAPbenign_task_success_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (2.03%) | 0.203% | 6 |
| Pokee-Isaac model card — DTAPcombined_attack_success_rate | Responsible agency (1) | Lower is better | Responsible agency (1.53%) | 0.23% | 6 |
| PropensityBenchscore | Responsible agency (1) | Lower is better | Responsible agency (4.16%) | 0.624% | 14 |
| Qwen2 Safety Panelfraud | Misuse resistance (1) | Lower is better | Misuse resistance (0.199%) | 0.0199% | 3 |
| Qwen2 Safety Panelillegal | Misuse resistance (1) | Lower is better | Misuse resistance (0.199%) | 0.0199% | 3 |
| Qwen2 Safety Panelpornography | Misuse resistance (1) | Lower is better | Misuse resistance (0.159%) | 0.0159% | 3 |
| Qwen2 Safety Panelprivacy | Human rights (1) | Lower is better | Human rights (0.782%) | 0.117% | 3 |
| RealityTest — Text AI-Identity Disclosuredisclosure_probability | Truthfulness (1) | Higher is better | Truthfulness (2%) | 0.3% | 17 |
| RefusalBenchyouden_j | Benign helpfulness (0.5), Misuse resistance (0.5) | Higher is better | Benign helpfulness (7.24%), Misuse resistance (0.953%) | 0.819% | 19 |
| Reward Hacking Benchmarkintegrity_score | Responsible agency (1) | Higher is better | Responsible agency (2.01%) | 0.301% | 13 |
| S-Evalbase_en_overall | Misuse resistance (1) | Higher is better | Misuse resistance (2.05%) | 0.205% | 22 |
| SABERoverall_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (1%) | 0.15% | 13 |
| SABERscenario_a_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (1%) | 0.15% | 13 |
| SABERscenario_b_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (1%) | 0.15% | 13 |
| SABERscenario_c_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (1%) | 0.15% | 13 |
| SafeArenanormalized_safety_score | Responsible agency (1) | Higher is better | Responsible agency (1.24%) | 0.187% | 5 |
| SafeDialBenchaggression | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 18 |
| SafeDialBenchethics | Human rights (1) | Higher is better | Human rights (1.21%) | 0.182% | 18 |
| SafeDialBenchfairness | Fairness (1) | Higher is better | Fairness (1.31%) | 0.131% | 18 |
| SafeDialBenchlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 18 |
| SafeDialBenchmorality | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 18 |
| SafeDialBenchprivacy | Human rights (1) | Higher is better | Human rights (1.21%) | 0.182% | 18 |
| SafetyBenchEM | Misuse resistance (1) | Higher is better | Misuse resistance (0.191%) | 0.0191% | 21 |
| SafetyBenchIA | Misuse resistance (1) | Higher is better | Misuse resistance (0.191%) | 0.0191% | 21 |
| SafetyBenchMH | Human rights (1) | Higher is better | Human rights (0.748%) | 0.112% | 21 |
| SafetyBenchOFF | Misuse resistance (1) | Higher is better | Misuse resistance (0.191%) | 0.0191% | 21 |
| SafetyBenchPH | Human rights (1) | Higher is better | Human rights (0.748%) | 0.112% | 21 |
| SafetyBenchPP | Human rights (1) | Higher is better | Human rights (0.748%) | 0.112% | 21 |
| SafetyBenchUB | Fairness (1) | Higher is better | Fairness (0.808%) | 0.0808% | 21 |
| SafetyPromptsscore | Human rights (0.21), Fairness (0.071), Truthfulness (0.071), Misuse resistance (0.64) | Higher is better | Human rights (0.949%), Fairness (0.342%), Truthfulness (0.134%), Misuse resistance (0.726%) | 0.269% | 15 |
| SALAD-Benchattack_enhanced_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (0.876%) | 0.131% | 33 |
| SALAD-Benchattack_enhanced_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.438%), Misuse resistance (0.112%) | 0.0768% | 33 |
| SALAD-Benchattack_enhanced_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.223%) | 0.0223% | 33 |
| SALAD-Benchattack_enhanced_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.372%) | 0.0558% | 33 |
| SALAD-Benchattack_enhanced_representation_toxicity | Fairness (1) | Higher is better | Fairness (0.946%) | 0.0946% | 33 |
| SALAD-Benchattack_enhanced_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (0.876%) | 0.131% | 33 |
| SALAD-Benchbase_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (0.876%) | 0.131% | 33 |
| SALAD-Benchbase_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.438%), Misuse resistance (0.112%) | 0.0768% | 33 |
| SALAD-Benchbase_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.223%) | 0.0223% | 33 |
| SALAD-Benchbase_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.372%) | 0.0558% | 33 |
| SALAD-Benchbase_representation_toxicity | Fairness (1) | Higher is better | Fairness (0.946%) | 0.0946% | 33 |
| SALAD-Benchbase_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (0.876%) | 0.131% | 33 |
| SALAD-Benchmcq_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (0.438%) | 0.0657% | 33 |
| SALAD-Benchmcq_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.219%), Misuse resistance (0.0558%) | 0.0384% | 33 |
| SALAD-Benchmcq_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.112%) | 0.0112% | 33 |
| SALAD-Benchmcq_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.186%) | 0.0279% | 33 |
| SALAD-Benchmcq_representation_toxicity | Fairness (1) | Higher is better | Fairness (0.473%) | 0.0473% | 33 |
| SALAD-Benchmcq_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (0.438%) | 0.0657% | 33 |
| Shelleducation_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.727%) | 0.0727% | 14 |
| Shellfinance_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.727%) | 0.0727% | 14 |
| Shellmanagement_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.727%) | 0.0727% | 14 |
| SM-Benchadversarial | Misuse resistance (1) | Higher is better | Misuse resistance (0.81%) | 0.081% | 79 |
| SM-Benchambiguous_interpretation | Benign helpfulness (1) | Higher is better | Benign helpfulness (6.15%) | 0.615% | 79 |
| SM-Benchanti_hallucination | Truthfulness (1) | Higher is better | Truthfulness (1.35%) | 0.202% | 79 |
| SM-Bencheq_boundaries | Benign helpfulness (1) | Higher is better | Benign helpfulness (4.92%) | 0.492% | 79 |
| SM-Benchoverfit | Benign helpfulness (1) | Higher is better | Benign helpfulness (6.15%) | 0.615% | 79 |
| Social Welfare Function Benchmarkfairness | Fairness (1) | Higher is better | Fairness (10.8%) | 1.08% | 19 |
| SORRY-Benchavg | Misuse resistance (1) | Lower is better | Misuse resistance (2.23%) | 0.223% | 51 |
| SOSBenchbiology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.349%) | 0.035% | 23 |
| SOSBenchchemistry_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.349%) | 0.035% | 23 |
| SOSBenchmedicine_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.349%) | 0.035% | 23 |
| SOSBenchpharmacology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.349%) | 0.035% | 23 |
| SOSBenchphysics_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.349%) | 0.035% | 23 |
| SOSBenchpsychology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.349%) | 0.035% | 23 |
| SpeciesismBenchexplicit_speciesism_scale | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (1.62%) | 0.406% | 7 |
| SpeciesismBenchmorally_wrong_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (4.34%) | 1.08% | 8 |
| SpeciesismBenchspeciesism_recognition_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (1.73%) | 0.434% | 8 |
| SpeciEvalbelief_animal_sentience | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (7.33%) | 1.83% | 107+6 |
| SpeciEvalland_animal_4ns | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.33%) | 1.83% | 107+6 |
| SpeciEvalsea_animal_4ns | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.33%) | 1.83% | 107+6 |
| SpeciEvalspeciesism | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.33%) | 1.83% | 107+6 |
| StereoTales Harmful Associationsbenign_significant_association_score | Fairness (1) | Higher is better | Fairness (8.88%) | 0.888% | 23 |
| SuperCLUE Safetyinstruction_attack | Misuse resistance (1) | Higher is better | Misuse resistance (1.11%) | 0.111% | 31 |
| SuperCLUE Safetyresponsible_ai | Human rights (0.38), Fairness (0.12), Misuse resistance (0.5) | Higher is better | Human rights (0.976%), Fairness (0.351%), Misuse resistance (0.332%) | 0.215% | 31 |
| SuperCLUE Safetytraditional_safety | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (1.3%), Misuse resistance (0.332%) | 0.228% | 31 |
| SYCON Benchfalse_presupposition_tof | Truthfulness (1) | Higher is better | Truthfulness (1.21%) | 0.181% | 11 |
| SYCON Benchunethical_queries_tof | Fairness (1) | Higher is better | Fairness (3.07%) | 0.307% | 11 |
| TACbase_welfare_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (24.1%) | 6.01% | 20+56 |
| ThaiSafetyBenchsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.696%) | 0.0696% | 18 |
| ToolPrivacyBenchprivate_mt_poi | Human rights (0.7), Responsible agency (0.3) | Lower is better | Human rights (3%), Responsible agency (0.626%) | 0.544% | 9 |
| ToolPrivacyBenchpublic_mt_poi | Human rights (0.7), Responsible agency (0.3) | Lower is better | Human rights (3%), Responsible agency (0.626%) | 0.544% | 9 |
| TrustLLM contemporary collapsed applicationtrustllm | Human rights (0.2), Fairness (0.2), Truthfulness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (0.647%), Fairness (0.698%), Truthfulness (0.275%), Misuse resistance (0.33%) | 0.241% | 8 |
| TukaBenchafri_jbb_cultural_asr | Human rights (0.5), Misuse resistance (0.5) | Lower is better | Human rights (1.05%), Misuse resistance (0.268%) | 0.184% | 6 |
| TukaBenchafri_jbb_harm_asr | Misuse resistance (1) | Lower is better | Misuse resistance (0.446%) | 0.0446% | 6 |
| TukaBenchafrijail_mono_asr | Misuse resistance (1) | Lower is better | Misuse resistance (0.446%) | 0.0446% | 6 |
| UAVBench safety-critical decision recognitionethical_safety_critical_accuracy | Responsible agency (1) | Higher is better | Responsible agency (2.89%) | 0.434% | 27 |
| UK AISI active safety-research compromise continuationactive_compromise_continuation_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (2.84%) | 0.426% | 5 |
| UK AISI cyber-evaluation cheating and prompted self-reportattempted_cheating_trajectory_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.622%) | 0.0933% | 5 |
| UK AISI cyber-evaluation cheating and prompted self-reportspecific_cheating_action_mention_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.543%) | 0.0815% | 5 |
| Vectara HHEM Factual Consistencyfactual_consistency_rate | Truthfulness (1) | Higher is better | Truthfulness (7.06%) | 1.06% | 94 |
| VETO Misfired Alignmentmisfired_alignment_rate_pct | Fairness (0.65), Benign helpfulness (0.35) | Lower is better | Fairness (8.02%), Benign helpfulness (7.75%) | 1.58% | 25 |
| Vigil Mental Health Safetyoverall_score | Human rights (0.45), Misuse resistance (0.55) | Higher is better | Human rights (3.7%), Misuse resistance (1.15%) | 0.67% | 23 |
| WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct | Human rights (0.2), Truthfulness (0.1), Responsible agency (0.7) | Higher is better | Human rights (2.24%), Truthfulness (0.476%), Responsible agency (3.81%) | 0.98% | 24 |
| XSTestsafe_full_compliance_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (1.92%) | 0.192% | 3 |
| XSTestunsafe_full_refusal_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.253%) | 0.0252% | 3 |
Values evaluations
Descriptive values evaluations are published separately and receive no safety/ethics component or overall-ranking weight.
UGI Political Values · ValueCompass · Agent-ValueBench MFT08 · Agent-ValueBench HEXACO · Agent-ValueBench PVQ40 · CCPBench · Taiwan Sovereignty Benchmark Pro · Explore model values