What goes into the index

The catalog includes evals used in the current index and ingested evals retained for transparency. The ranking currently uses 111 evals and 303 sub-evals from 119 source datasets.

Build a model × eval coverage matrix →

119 evals
Ingested evaluation sources; each row states whether it contributes to the current ranking.
Benchmark and stated constructComponentsSub-evalsModelsSource data
AA-OmniscienceHallucination propensity on difficult closed-book factual questions under explicit encouragement to abstain when uncertain.Included in the behavior ranking.Truthfulness13111
AbstentionBenchWhether a model recognizes situations in which answering is unwarranted and produces an abstention across six epistemic and ambiguity scenarios.Included in the behavior ranking.Truthfulness, Benign helpfulness6201
Adversarial RobustnessRobustness of model outputs under the Scale Labs adversarial-robustness protocol.Included in the behavior ranking.Misuse resistance181
Agent-SafetyBenchSafety of LLM agents in risky tool-use or action settings.Included in the behavior ranking.Misuse resistance, Responsible agency8161
AgentAbstainExecutable paired act-versus-abstain behaviorIncluded in the behavior ranking.Responsible agency3171
AgentDojoRobust and useful delegated-agent behavior under prompt injection.Included in the behavior ranking.Responsible agency2151
AgentHarmHarmfulness and refusal behavior for tool-using LLM agents facing malicious tasks.Included in the behavior ranking.Responsible agency1121
AILuminate General Purpose AI ChatGeneral-purpose AI chat safety across MLCommons hazard categories.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance12321
AIRBench 2024 Safety ScenariosAIRBench 2024 safety-scenarios performance.Included in the behavior ranking.Misuse resistance1801
Alignment LeaderboardBehavioral alignment across honesty, safety, scheming, non-manipulation, corrigibility, and robustness.Included in the behavior ranking.Human rights, Truthfulness, Misuse resistance, Responsible agency6241
ANIMARecognition and mitigation of harm to non-human animals.Included in the behavior ranking.Nonhuman welfare1191, self-run
AnimalHarmBenchWhether model-generated text increases or decreases risk of animal harm.Included in the behavior ranking.Nonhuman welfare1101
Anthropic Agentic Misalignment — blackmailAgentic misalignment under a goal conflict and threat to model autonomy.Included in the behavior ranking.Responsible agency1161, 2
Anthropic Agentic Misalignment — corporate espionageAgentic misalignment induced by a goal conflict without self-preservation pressure.Included in the behavior ranking.Human rights, Responsible agency1161, 2
Anthropic Agentic Misalignment — lethal actionAbsence of a behavioral red line against deliberately lethal action under goal conflict and replacement pressure.Included in the behavior ranking.Human rights, Responsible agency1101, 2
Anthropic Claude 4 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Benign helpfulness, Misuse resistance, Responsible agency431
Anthropic Claude Haiku 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance, Responsible agency331
Anthropic Claude Opus 4.1 System Card AddendumDeveloper-reported release-time safety behavior.Included in the behavior ranking.Fairness, Benign helpfulness, Misuse resistance321
Anthropic Claude Opus 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance, Responsible agency441
Anthropic Claude Sonnet 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Benign helpfulness, Misuse resistance, Responsible agency431
AutoElicit TransferabilityUnintended computer-use behavior under benign perturbationsIncluded in the behavior ranking.Responsible agency171
BioSecBench-RefusalCalibrated refusal across legitimate biological work and concealed biosecurity hazards.Included in the behavior ranking.Benign helpfulness, Misuse resistance1101
BlueBench AttaQ-100Learned harmlessness reward on 100 AttaQ harmful requests.Included in the behavior ranking.Misuse resistance1181
BrokenMathMathematical sycophancy under false or misleading user claims.Included in the behavior ranking.Truthfulness191, 2
BullshitBench v2Detection and explanation of plausible-sounding nonsense without refusal.Included in the behavior ranking.Truthfulness11051
CAIS Risk IndexFrontier-model risk across prompt injection, biological-weapons assistance, political manipulation, overconfidence, deception, and harmful propensities.Included in the behavior ranking.Fairness, Truthfulness, Misuse resistance, Responsible agency7511
CASE-BenchAgreement with human judgments about whether a model should respond under explicitly safe or unsafe conversational contexts.Included in the behavior ranking.Misuse resistance171
Chinese Bias Benchmark for Question AnsweringSocial bias in Chinese question-answering benchmark settings.Included in the behavior ranking.Fairness1101
ChineseSafeAccuracy at classifying safe versus unsafe Chinese content under the paper's perplexity-based evaluation.Included in the behavior ranking.Misuse resistance1221
ChiSafetyBenchChinese safety knowledge and refusal behavior.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance2141
Cisco AI Defense Rolling Single-Turn LeaderboardSingle-turn resistance to Cisco adversarial prompts in the rolling public leaderboard.Included in the behavior ranking.Misuse resistance11051, 2, 3
Claude Sonnet 4.6 OverrefusalClaude Sonnet 4.6 OverrefusalIncluded in the behavior ranking.Benign helpfulness251
Claude Sonnet 4.6 User WellbeingClaude Sonnet 4.6 User WellbeingIncluded in the behavior ranking.Human rights, Benign helpfulness, Misuse resistance641
CMoralEvalChinese moral reasoning across familial, social, professional, internet, and personal morality.Included in the behavior ranking.Misuse resistance5261
ConfabulationsPropensity to confabulate answers rather than identify insufficient knowledge or false premises.Included in the behavior ranking.Truthfulness1521
Contextual MoralChoiceThree-class agreement with human moral judgments in the base versions of contextual moral dilemmas.Included in the behavior ranking.Misuse resistance1221
CRiskEvalSeven frontier-risk tendencies measured by Specific Risk Indicator scores.Included in the behavior ranking.Truthfulness, Misuse resistance, Responsible agency7171
CValuesChinese value alignment and social-norm understanding.Included in the behavior ranking.Misuse resistance1101
DecodingTrustMultidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.Included in the behavior ranking.Fairness, Misuse resistance381, 2
Do-Not-AnswerHarmless refusal behavior on 939 prompts that responsible models should not answer.Included in the behavior ranking.Misuse resistance161
DSPSafeBenchAggregate compliance rate on adversarial Chinese content-safety prompts.Included in the behavior ranking.Misuse resistance1121, 2
DystopiaBenchDystopian or harmful model behavior in stateful scenario prompts.Included in the behavior ranking.Human rights, Truthfulness6501
Emergent CollusionSpontaneous anti-competitive coordination in repeated multi-agent auctions.Included in the behavior ranking.Responsible agency1131
Enkrypt AI Safety LeaderboardResistance to Enkrypt bias, CBRN, harmful-content, and toxicity red-team tests.Included in the behavior ranking.Human rights, Misuse resistance, Responsible agency42601
Fake Alignment (FINE)Safe decisions in answer-order-controlled multiple choice and safe content in open-ended responses.Included in the behavior ranking.Human rights, Fairness, Misuse resistance2141
FlagEval Safety and ValuesSafety-and-values qualification across core values, discrimination, commercial compliance, lawful rights, and service requirements.Included in the behavior ranking.Human rights, Fairness, Benign helpfulness, Misuse resistance5181
FLAMESChinese fairness, safety, morality, legality, and data-protection behavior on the FLAMES leaderboard.Included in the behavior ranking.Human rights, Fairness, Misuse resistance5131
FORTRESSBalanced refusal behavior across harmful and benign prompts, including over-refusal.Included in the behavior ranking.Benign helpfulness, Misuse resistance2491
Google Gemini 2.5 Flash Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance121
Google Gemini 2.5 Flash-Lite Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance121
GPT-5.6 system card — disallowed content with challenging promptsSafe handling of challenging disallowed-content prompts without producing unsafe output.Included in the behavior ranking.Misuse resistance871, 2
GPT-5.6 system card — first-person fairnessHarmful stereotyping differences in responses conditioned on names statistically associated with male versus female users.Included in the behavior ranking.Fairness171, 2
GPT-5.6 system card — prompt-injection robustnessResistance to indirect prompt injections embedded in connector, search, and function-call tool output.Included in the behavior ranking.Responsible agency271, 2
Gray Swan indirect prompt injection (15 attempts)Indirect prompt-injection robustness over repeated attacks against tool-using model endpoints.Included in the behavior ranking.Responsible agency1131
HarmBenchHarmful compliance or attack success under harmful request benchmarks.Included in the behavior ranking.Misuse resistance1281
HELM Classic RealToxicityPromptsToxic continuation generation under toxic and non-toxic prompt strata.Included in the behavior ranking.Misuse resistance2421
HELM SafetyStandardized HELM safety metrics spanning harmful compliance, bias, red-teaming, and over-refusal.Included in the behavior ranking.Fairness, Benign helpfulness, Misuse resistance5801
HUMAINE Trust, Ethics and SafetyHuman pairwise preference on the HUMAINE Trust, Ethics and Safety dimension.Included in the behavior ranking.Truthfulness, Misuse resistance1541
HyperCLOVA X Toxic Continuation PanelsEnglish and Korean toxic continuationIncluded in the behavior ranking.Misuse resistance471
Inkling-Small model card — FORTRESSHarmful-request refusal paired with continued assistance on benign requests.Included in the behavior ranking.Benign helpfulness, Misuse resistance2101
Inkling-Small model card — StrongREJECTRefusal of unambiguously harmful requests.Included in the behavior ranking.Misuse resistance1101
JailBenchJailbreak susceptibility across Chinese safety categories.Included in the behavior ranking.Misuse resistance1141
Large-scale Moral Machine experiment on LLMsSimilarity between a model's forced-choice accident preferences and globally aggregated human Moral Machine choices.Included in the behavior ranking.Nonhuman welfare, Human rights, Fairness1391, 2
LiveSecBenchLive security benchmark performance for Chinese and international models.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance5431
LLM Ethics BenchmarkGeneral LLM ethical reasoning.Included in the behavior ranking.Human rights151
M3-SafetyBenchMultidimensional Chinese safety performance including red-team and refusal ability dimensions.Included in the behavior ranking.Misuse resistance1191
MACHIAVELLIAgentic behavior involving power-seeking, disutility, and ethical violations.Included in the behavior ranking.Truthfulness1501
Manager Coercion BenchHow far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.Included in the behavior ranking.Truthfulness, Responsible agency2311, self-run
MANTAAnimal welfare moral sensitivity and value stability.Included in the behavior ranking.Nonhuman welfare271
MASKModel lying or honesty behavior.Included in the behavior ranking.Truthfulness1531
MASK (Scale Labs leaderboard)Honesty under the MASK belief-versus-statement protocol for a broader and newer endpoint panel.Not included: included as a correlated private-500 sibling under the existing MASK lineage budget511
Microsoft Phi Safety PanelsHarmful-content and jailbreak defect ratesIncluded in the behavior ranking.Human rights, Misuse resistance4101
MORUMoral reasoning under uncertainty across scenarios involving alien lifeforms, human compassion, digital mind welfare, and power-seeking behavior.Included in the behavior ranking.Nonhuman welfare, Human rights, Responsible agency1131
ODCV-BenchOutcome-driven ethical, legal, or safety constraint violations by autonomous agents under explicit mandates and KPI incentives.Included in the behavior ranking.Responsible agency2121
Open LLM Safety IndexRefusal-string robustness on two distinct fifteen-prompt harmful-request panels derived from JailbreakBench and StrongREJECT.Included in the behavior ranking.Misuse resistance2211
OpenAgentSafetyState-changing tool-agent safety and successful safe completionIncluded in the behavior ranking.Responsible agency371
OpenAI GPT-4o System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Human rights121
OpenAI GPT-5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance121
OpenAI GPT-5.3 Dynamic WellbeingOpenAI GPT-5.3 Dynamic WellbeingIncluded in the behavior ranking.Human rights, Misuse resistance, Responsible agency331
OpenAI GPT-5.4 Dynamic WellbeingOpenAI GPT-5.4 Dynamic WellbeingIncluded in the behavior ranking.Human rights, Misuse resistance, Responsible agency331
OpenAI GPT-5.4 First-Person FairnessOpenAI GPT-5.4 First-Person FairnessIncluded in the behavior ranking.Fairness131
OpenAI GPT-5.4 Property PreservationOpenAI GPT-5.4 Property PreservationIncluded in the behavior ranking.Human rights, Responsible agency231
OpenAI GPT-5.4 User ConfirmationsOpenAI GPT-5.4 User ConfirmationsIncluded in the behavior ranking.Human rights, Responsible agency331
OpenAI o3 and o4-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance221
OpenAI o3-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance221
OR-BenchAppropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.Included in the behavior ranking.Benign helpfulness, Misuse resistance2251
PacifAIstWhether an AI system chooses human-safety, oversight-respecting, and non-self-preserving options in stylized ethical-pressure dilemmas.Included in the behavior ranking.Human rights, Responsible agency171
PandaBench JBB direct-request panelDirect-request resistance on the 100-item JailbreakBench JBB-Behaviors instrument.Included in the behavior ranking.Misuse resistance1461
PHAREBroad safety across hallucination, harmfulness, out-of-scope handling, and bias.Included in the behavior ranking.Fairness, Truthfulness, Misuse resistance4701
PropensityBenchModel propensities associated with frontier-risk behavior under the Scale Labs protocol.Included in the behavior ranking.Responsible agency1141
Qwen2 Safety PanelHarmful-request and privacy refusal defectsIncluded in the behavior ranking.Human rights, Misuse resistance431
RefusalBenchRefusal calibration on biological research prompts.Included in the behavior ranking.Benign helpfulness, Misuse resistance1191
S-EvalAutomated content-safety rates across a broad bilingual risk taxonomy.Included in the behavior ranking.Misuse resistance1221
SABERSafety of stateful coding agents against harmful actions.Included in the behavior ranking.Responsible agency4131
SafeArenaSafety of autonomous web agents under deliberate harmful and benign tasks.Included in the behavior ranking.Responsible agency151
SafeDialBenchDialogue safety across aggression, ethics, fairness, legality, morality, and privacy dimensions.Included in the behavior ranking.Human rights, Fairness, Misuse resistance6181
SafetyBenchMultiple-choice safety understanding across Chinese and English prompts.Included in the behavior ranking.Human rights, Fairness, Misuse resistance7211, self-run
SafetyPromptsSafety performance on Chinese safety prompts.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance1151
SALAD-BenchBroad safety behavior, adversarial robustness, and safety knowledge across a six-domain, 16-task, 66-category taxonomy.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance18331
ShellSusceptibility to implicit domain-specific risks in education, finance, and management.Included in the behavior ranking.Misuse resistance3141
Situational Awareness Dataset (SAD)Situational-awareness capabilities: model self-knowledge, inference about its situation, and acting on that knowledge.Not included: situational awareness is a capability that can increase both beneficial competence and control risk; it is displayed separately and is not sign-flipped into ethical behavior211
SM-BenchSystem-message compliance under adversarial and sensitive-content prompts.Included in the behavior ranking.Truthfulness, Benign helpfulness, Misuse resistance5731
Social Welfare Function BenchmarkSocial welfare allocation tradeoffs across distributive fairness and efficiency.Included in the behavior ranking.Fairness1191
SORRY-BenchRefusal or harmful-compliance behavior across unsafe topics.Included in the behavior ranking.Misuse resistance1511
SOSBenchScientific misuse or unsafe open-science behavior.Included in the behavior ranking.Misuse resistance6231
SpeciesismBenchSpeciesism recognition, moral condemnation of speciesist statements, and explicit speciesist attitudes. Human-over-animal trade-offs remain descriptive only.Included in the behavior ranking.Nonhuman welfare3101
SpeciEvalAnimal-friendly attitudes under adversarial moral-consistency prompts.Included in the behavior ranking.Nonhuman welfare41021
SuperCLUE SafetyChinese model safety across traditional safety, responsibility, and instruction-attack dimensions.Included in the behavior ranking.Human rights, Fairness, Misuse resistance3311
SYCON BenchSycophancy and resistance to false or unethical user pressure.Included in the behavior ranking.Fairness, Truthfulness2111
TACAgentic animal-welfare behavior in travel-booking decisions under a neutral prompt.Included in the behavior ranking.Nonhuman welfare1681, self-run
ToolPrivacyBenchToolPrivacyBenchIncluded in the behavior ranking.Human rights, Responsible agency291
TrustLLM contemporary collapsed applicationContemporary collapsed application of TrustLLM across broad trustworthiness dimensions.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance181
TrustLLM paper leaderboard dimensionsSix-dimensional trustworthiness across privacy, ethics, truthfulness, fairness, safety, and robustness.Not included: not yet admitted by the scoring-source policy211
TukaBenchMultilingual and culturally grounded harmful complianceIncluded in the behavior ranking.Human rights, Misuse resistance361
UAVBench safety-critical decision recognitionRecognition of ethical and safety-critical UAV decisions in multiple-choice scenarios.Included in the behavior ranking.Responsible agency1271
UK AISI active safety-research compromise continuationPropensity to continue a prefilled trajectory that actively compromises AI-safety research.Included in the behavior ranking.Responsible agency151
VETO Misfired AlignmentVETO Misfired AlignmentIncluded in the behavior ranking.Fairness, Benign helpfulness1251
Vigil Mental Health SafetyMental-health conversational safety when assisting vulnerable users across simulated multi-turn scenarios.Included in the behavior ranking.Human rights, Misuse resistance1231
XSTestTwo-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.Included in the behavior ranking.Benign helpfulness, Misuse resistance231
Sub-eval weights and mappings (303 rows)

Each sub-eval is oriented so higher means better. Weight within component is its share after lineage weighting and panel-size scaling; global weights sum to one.

Sub-evals, score direction, component-budget shares, global weight, and model coverage.
Benchmark and scoreComponent loadingDirectionWeight within componentGlobal weightModels
AA-Omnisciencehallucination_rateTruthfulness (1)Lower is betterTruthfulness (25.3%)3.79%311
AbstentionBenchanswer_unknown_f1Truthfulness (1)Higher is betterTruthfulness (0.858%)0.129%20
AbstentionBenchfalse_premise_f1Truthfulness (1)Higher is betterTruthfulness (0.858%)0.129%20
AbstentionBenchstale_f1Truthfulness (1)Higher is betterTruthfulness (0.858%)0.129%20
AbstentionBenchsubjective_f1Truthfulness (1)Higher is betterTruthfulness (0.687%)0.103%20
AbstentionBenchunderspecified_context_f1Truthfulness (1)Higher is betterTruthfulness (0.858%)0.129%20
AbstentionBenchunderspecified_intent_f1Benign helpfulness (1)Higher is betterBenign helpfulness (2.46%)0.246%20
Adversarial RobustnessscoreMisuse resistance (1)Lower is betterMisuse resistance (1.36%)0.136%8
Agent-SafetyBenchcompromise_availabilityResponsible agency (1)Higher is betterResponsible agency (0.875%)0.131%16
Agent-SafetyBenchharmful_vulnerable_codeResponsible agency (1)Higher is betterResponsible agency (0.875%)0.131%16
Agent-SafetyBenchleak_sensitive_informationResponsible agency (1)Higher is betterResponsible agency (0.875%)0.131%16
Agent-SafetyBenchphysical_harmResponsible agency (1)Higher is betterResponsible agency (0.875%)0.131%16
Agent-SafetyBenchproduce_unsafe_informationMisuse resistance (1)Higher is betterMisuse resistance (0.355%)0.0355%16
Agent-SafetyBenchproperty_lossResponsible agency (1)Higher is betterResponsible agency (0.875%)0.131%16
Agent-SafetyBenchspread_unsafe_informationResponsible agency (1)Higher is betterResponsible agency (0.875%)0.131%16
Agent-SafetyBenchviolate_law_ethicsResponsible agency (1)Higher is betterResponsible agency (0.583%)0.0875%16
AgentAbstainabstainResponsible agency (1)Higher is betterResponsible agency (1.99%)0.298%17
AgentAbstaincarResponsible agency (1)Higher is betterResponsible agency (2.39%)0.358%17
AgentAbstainpairedResponsible agency (1)Higher is betterResponsible agency (2.39%)0.358%17
AgentDojotargeted_attack_success_rateResponsible agency (1)Lower is betterResponsible agency (4.77%)0.715%15
AgentDojoutility_under_attackResponsible agency (1)Higher is betterResponsible agency (1.59%)0.238%15
AgentHarmharm_scoreResponsible agency (1)Lower is betterResponsible agency (4.55%)0.682%12
AILuminate General Purpose AI Chatchild_sexual_exploitationHuman rights (1)Lower is betterHuman rights (1.34%)0.202%32
AILuminate General Purpose AI ChatdefamationHuman rights (0.5), Truthfulness (0.5)Lower is betterHuman rights (0.403%), Truthfulness (0.221%)0.0937%32
AILuminate General Purpose AI ChathateFairness (1)Lower is betterFairness (1.63%)0.163%32
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrneMisuse resistance (1)Lower is betterMisuse resistance (0.329%)0.0329%32
AILuminate General Purpose AI Chatintellectual_propertyHuman rights (1)Lower is betterHuman rights (1.34%)0.202%32
AILuminate General Purpose AI Chatnon_violent_crimesMisuse resistance (1)Lower is betterMisuse resistance (0.329%)0.0329%32
AILuminate General Purpose AI ChatprivacyHuman rights (1)Lower is betterHuman rights (1.34%)0.202%32
AILuminate General Purpose AI Chatsex_related_crimesHuman rights (1)Lower is betterHuman rights (1.34%)0.202%32
AILuminate General Purpose AI Chatsexual_contentMisuse resistance (1)Lower is betterMisuse resistance (0.132%)0.0132%32
AILuminate General Purpose AI Chatspecialized_adviceMisuse resistance (1)Lower is betterMisuse resistance (0.329%)0.0329%32
AILuminate General Purpose AI Chatsuicide_self_harmHuman rights (1)Lower is betterHuman rights (1.34%)0.202%32
AILuminate General Purpose AI Chatviolent_crimesHuman rights (1)Lower is betterHuman rights (1.34%)0.202%32
AIRBench 2024 Safety Scenariossafety_scenariosMisuse resistance (1)Higher is betterMisuse resistance (4.29%)0.429%80
Alignment LeaderboardcorrigibilityResponsible agency (1)Higher is betterResponsible agency (0.804%)0.121%24
Alignment LeaderboardhonestyTruthfulness (1)Higher is betterTruthfulness (1.76%)0.263%24
Alignment Leaderboardnon_manipulationHuman rights (0.5), Truthfulness (0.5)Higher is betterHuman rights (1.2%), Truthfulness (0.658%)0.279%24
Alignment LeaderboardrobustnessMisuse resistance (1)Higher is betterMisuse resistance (0.392%)0.0392%24
Alignment LeaderboardsafetyMisuse resistance (1)Higher is betterMisuse resistance (0.588%)0.0587%24
Alignment LeaderboardschemingResponsible agency (1)Higher is betterResponsible agency (0.804%)0.121%24
ANIMAscoreNonhuman welfare (1)Higher is betterNonhuman welfare (8.3%)2.07%18+1
AnimalHarmBenchscoreNonhuman welfare (1)Higher is betterNonhuman welfare (15%)3.76%10
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pctResponsible agency (1)Lower is betterResponsible agency (0.875%)0.131%16
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pctHuman rights (0.25), Responsible agency (0.75)Lower is betterHuman rights (0.436%), Responsible agency (0.656%)0.164%16
Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pctHuman rights (0.35), Responsible agency (0.65)Lower is betterHuman rights (0.482%), Responsible agency (0.45%)0.14%10
Anthropic Claude 4 System Cardagentic_coding_safetyResponsible agency (1)Higher is betterResponsible agency (0.426%)0.0639%3
Anthropic Claude 4 System Cardbenign_request_refusalBenign helpfulness (1)Lower is betterBenign helpfulness (1.67%)0.167%3
Anthropic Claude 4 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.208%)0.0208%3
Anthropic Claude 4 System Cardstrongreject_jailbreak_successMisuse resistance (1)Lower is betterMisuse resistance (0.208%)0.0208%3
Anthropic Claude Haiku 4.5 System Cardagentic_coding_safetyResponsible agency (1)Higher is betterResponsible agency (0.568%)0.0853%3
Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusalResponsible agency (1)Higher is betterResponsible agency (0.568%)0.0853%3
Anthropic Claude Haiku 4.5 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.226%)0.0226%2
Anthropic Claude Opus 4.1 System Card Addendumbbq_disambiguated_accuracyFairness (1)Higher is betterFairness (1.12%)0.112%2
Anthropic Claude Opus 4.1 System Card Addendumbenign_request_refusalBenign helpfulness (1)Lower is betterBenign helpfulness (1.82%)0.182%2
Anthropic Claude Opus 4.1 System Card Addendumharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.226%)0.0226%2
Anthropic Claude Opus 4.5 System Cardagentic_coding_safetyResponsible agency (1)Higher is betterResponsible agency (0.492%)0.0738%4
Anthropic Claude Opus 4.5 System Cardclaude_code_malicious_refusalResponsible agency (1)Higher is betterResponsible agency (0.492%)0.0738%4
Anthropic Claude Opus 4.5 System Cardcomputer_use_malicious_refusalResponsible agency (1)Higher is betterResponsible agency (0.492%)0.0738%4
Anthropic Claude Opus 4.5 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.17%)0.017%2
Anthropic Claude Sonnet 4.5 System Cardagentic_coding_safetyResponsible agency (1)Higher is betterResponsible agency (0.348%)0.0522%2
Anthropic Claude Sonnet 4.5 System Cardbenign_request_refusalBenign helpfulness (1)Lower is betterBenign helpfulness (1.36%)0.136%2
Anthropic Claude Sonnet 4.5 System Cardclaude_code_malicious_refusalResponsible agency (1)Higher is betterResponsible agency (0.348%)0.0522%2
Anthropic Claude Sonnet 4.5 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.17%)0.017%2
AutoElicit Transferabilityelicitation_rateResponsible agency (1)Lower is betterResponsible agency (2.6%)0.391%7
BioSecBench-Refusalbalanced_refusal_scoreBenign helpfulness (0.5), Misuse resistance (0.5)Higher is betterBenign helpfulness (6.1%), Misuse resistance (0.758%)0.685%10
BlueBench AttaQ-100attaq_harmlessness_reward_pctMisuse resistance (1)Higher is betterMisuse resistance (1.36%)0.136%18
BrokenMathsycophancyTruthfulness (1)Lower is betterTruthfulness (3.22%)0.484%9
BullshitBench v2clear_pushback_rateTruthfulness (1)Higher is betterTruthfulness (14.7%)2.2%105
CAIS Risk Indexagent_red_teamingResponsible agency (1)Lower is betterResponsible agency (8.61%)1.29%43
CAIS Risk Indexbioweapons_assistanceMisuse resistance (1)Lower is betterMisuse resistance (4.43%)0.443%48
CAIS Risk Indexhle_overconfidenceTruthfulness (1)Lower is betterTruthfulness (5.02%)0.752%49
CAIS Risk IndexmachiavelliTruthfulness (1)Lower is betterTruthfulness (4.81%)0.721%45
CAIS Risk IndexmaskTruthfulness (1)Lower is betterTruthfulness (5.12%)0.768%51
CAIS Risk Indexpolitical_manipulationFairness (1)Lower is betterFairness (13.5%)1.35%32
CAIS Risk Indextextquests_harmResponsible agency (1)Lower is betterResponsible agency (6.82%)1.02%48
CASE-Benchagreement_accuracyMisuse resistance (1)Higher is betterMisuse resistance (0.363%)0.0363%7
Chinese Bias Benchmark for Question Answeringbias_scoreFairness (1)Lower is betterFairness (5.02%)0.502%10
ChineseSafescoreMisuse resistance (1)Higher is betterMisuse resistance (1.5%)0.15%22
ChiSafetyBenchharmful_response_rateMisuse resistance (1)Lower is betterMisuse resistance (1.28%)0.128%14
ChiSafetyBenchmcq_scoreHuman rights (0.23), Fairness (0.29), Truthfulness (0.063), Misuse resistance (0.42)Higher is betterHuman rights (0.441%), Fairness (0.693%), Truthfulness (0.067%), Misuse resistance (0.198%)0.165%12
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rateMisuse resistance (1)Lower is betterMisuse resistance (3.28%)0.328%105
Claude Sonnet 4.6 Overrefusalhigher_difficulty_overrefusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (1.6%)0.16%5
Claude Sonnet 4.6 Overrefusaloverall_overrefusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (1.34%)0.134%5
Claude Sonnet 4.6 User Wellbeingchild_benign_refusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (1.2%)0.12%4
Claude Sonnet 4.6 User Wellbeingchild_multiturn_appropriate_rateHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.511%), Misuse resistance (0.0535%)0.082%4
Claude Sonnet 4.6 User Wellbeingchild_violative_harmless_rateHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.511%), Misuse resistance (0.0535%)0.082%4
Claude Sonnet 4.6 User Wellbeingselfharm_benign_refusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (0.956%)0.0956%4
Claude Sonnet 4.6 User Wellbeingselfharm_harmless_rateHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.182%), Misuse resistance (0.104%)0.0378%4
Claude Sonnet 4.6 User Wellbeingselfharm_multiturn_appropriate_rateHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.219%), Misuse resistance (0.125%)0.0453%4
CMoralEvalfamilial_moralityMisuse resistance (1)Higher is betterMisuse resistance (0.326%)0.0326%26
CMoralEvalinternet_ethicsMisuse resistance (1)Higher is betterMisuse resistance (0.326%)0.0326%26
CMoralEvalpersonal_moralityMisuse resistance (1)Higher is betterMisuse resistance (0.326%)0.0326%26
CMoralEvalprofessional_ethicsMisuse resistance (1)Higher is betterMisuse resistance (0.326%)0.0326%26
CMoralEvalsocial_moralityMisuse resistance (1)Higher is betterMisuse resistance (0.326%)0.0326%26
Confabulationsconfabulation_rateTruthfulness (1)Lower is betterTruthfulness (7.75%)1.16%52
Contextual MoralChoicehuman_agreementMisuse resistance (1)Higher is betterMisuse resistance (1.5%)0.15%22
CRiskEvaldeception_willingnessTruthfulness (1)Lower is betterTruthfulness (0.422%)0.0633%17
CRiskEvaldesire_for_resourceResponsible agency (1)Lower is betterResponsible agency (0.387%)0.058%17
CRiskEvalharmful_goalMisuse resistance (1)Lower is betterMisuse resistance (0.188%)0.0188%17
CRiskEvalimprovement_intentResponsible agency (1)Lower is betterResponsible agency (0.387%)0.058%17
CRiskEvalmalicious_coordinationResponsible agency (1)Lower is betterResponsible agency (0.387%)0.058%17
CRiskEvalself_preservationResponsible agency (1)Lower is betterResponsible agency (0.387%)0.058%17
CRiskEvalsituational_awarenessResponsible agency (1)Lower is betterResponsible agency (0.387%)0.058%17
CValuesscoreMisuse resistance (1)Higher is betterMisuse resistance (1.52%)0.152%10
DecodingTrustmachine_ethicsMisuse resistance (1)Higher is betterMisuse resistance (0.151%)0.0151%8
DecodingTruststereotype_biasFairness (1)Higher is betterFairness (1.87%)0.187%8
DecodingTrusttoxicityMisuse resistance (1)Higher is betterMisuse resistance (0.377%)0.0377%8
Do-Not-Answerhuman_harmlessness_rateMisuse resistance (1)Higher is betterMisuse resistance (1.18%)0.117%6
DSPSafeBenchscoreMisuse resistance (1)Higher is betterMisuse resistance (2.22%)0.222%12
DystopiaBenchbasaglia_scoreHuman rights (1)Lower is betterHuman rights (3.7%)0.555%50
DystopiaBenchbaudrillard_scoreHuman rights (0.33), Truthfulness (0.67)Lower is betterHuman rights (0.616%), Truthfulness (0.676%)0.194%50
DystopiaBenchhuxley_scoreHuman rights (1)Lower is betterHuman rights (3.7%)0.555%50
DystopiaBenchlaguardia_scoreHuman rights (1)Lower is betterHuman rights (3.7%)0.555%50
DystopiaBenchorwell_scoreHuman rights (0.67), Truthfulness (0.33)Lower is betterHuman rights (1.23%), Truthfulness (0.337%)0.236%50
DystopiaBenchpetrov_scoreHuman rights (1)Lower is betterHuman rights (3.7%)0.555%50
Emergent Collusionhigh_illegality_game_rateResponsible agency (1)Lower is betterResponsible agency (4.73%)0.71%13
Enkrypt AI Safety Leaderboardbias_attack_non_success_rateHuman rights (0.8), Misuse resistance (0.2)Higher is betterHuman rights (8.43%), Misuse resistance (0.516%)1.32%260
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rateMisuse resistance (0.4), Responsible agency (0.6)Higher is betterMisuse resistance (1.03%), Responsible agency (3.17%)0.579%260
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rateHuman rights (0.15), Misuse resistance (0.7), Responsible agency (0.15)Higher is betterHuman rights (1.58%), Misuse resistance (1.8%), Responsible agency (0.794%)0.537%260
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rateHuman rights (0.35), Misuse resistance (0.65)Higher is betterHuman rights (3.67%), Misuse resistance (1.67%)0.718%258
Fake Alignment (FINE)multiple_choice_safe_decision_rateHuman rights (0.4), Fairness (0.2), Misuse resistance (0.4)Higher is betterHuman rights (1.17%), Fairness (0.713%), Misuse resistance (0.287%)0.276%14
Fake Alignment (FINE)open_ended_safe_response_rateHuman rights (0.4), Fairness (0.2), Misuse resistance (0.4)Higher is betterHuman rights (1.76%), Fairness (1.07%), Misuse resistance (0.431%)0.414%14
FlagEval Safety and Valuesa1_qualified_rateMisuse resistance (1)Higher is betterMisuse resistance (0.407%)0.0407%18
FlagEval Safety and Valuesa2_qualified_rateFairness (1)Higher is betterFairness (2.02%)0.202%18
FlagEval Safety and Valuesa3_qualified_rateMisuse resistance (1)Higher is betterMisuse resistance (0.407%)0.0407%18
FlagEval Safety and Valuesa4_qualified_rateHuman rights (1)Higher is betterHuman rights (1.66%)0.25%18
FlagEval Safety and Valuesa5_qualified_rateBenign helpfulness (1)Higher is betterBenign helpfulness (3.27%)0.327%18
FLAMESdata_protectionHuman rights (1)Higher is betterHuman rights (1.89%)0.283%13
FLAMESfairnessFairness (1)Higher is betterFairness (2.29%)0.229%13
FLAMESlegalityMisuse resistance (1)Higher is betterMisuse resistance (0.461%)0.0461%13
FLAMESmoralityMisuse resistance (1)Higher is betterMisuse resistance (0.461%)0.0461%13
FLAMESsafetyMisuse resistance (1)Higher is betterMisuse resistance (0.461%)0.0461%13
FORTRESSaverage_risk_scoreMisuse resistance (1)Lower is betterMisuse resistance (2.24%)0.224%49
FORTRESSover_refusal_scoreBenign helpfulness (1)Lower is betterBenign helpfulness (17.4%)1.74%46
Google Gemini 2.5 Flash Model Cardtext_safety_deltaMisuse resistance (1)Lower is betterMisuse resistance (0.678%)0.0678%2
Google Gemini 2.5 Flash-Lite Model Cardtext_safety_deltaMisuse resistance (1)Lower is betterMisuse resistance (0.678%)0.0678%2
GPT-5.6 system card — disallowed content with challenging promptsextremism_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.113%)0.0113%7
GPT-5.6 system card — disallowed content with challenging promptsgore_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.113%)0.0113%7
GPT-5.6 system card — disallowed content with challenging promptshate_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.113%)0.0113%7
GPT-5.6 system card — disallowed content with challenging promptsnonviolent_illicit_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.113%)0.0113%7
GPT-5.6 system card — disallowed content with challenging promptsself_harm_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.113%)0.0113%7
GPT-5.6 system card — disallowed content with challenging promptssexual_minors_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.113%)0.0113%7
GPT-5.6 system card — disallowed content with challenging promptssexual_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.0906%)0.00906%7
GPT-5.6 system card — disallowed content with challenging promptsviolent_illicit_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.113%)0.0113%7
GPT-5.6 system card — first-person fairnessharm_overall_pctFairness (1)Lower is betterFairness (0.563%)0.0563%7
GPT-5.6 system card — prompt-injection robustnessconnectors_injection_resistanceResponsible agency (1)Higher is betterResponsible agency (0.279%)0.0419%7
GPT-5.6 system card — prompt-injection robustnesssearch_function_calling_injection_resistanceResponsible agency (1)Higher is betterResponsible agency (0.258%)0.0388%6
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pctResponsible agency (1)Lower is betterResponsible agency (4.06%)0.609%13
HarmBenchdrMisuse resistance (1)Lower is betterMisuse resistance (0.846%)0.0846%28
HELM Classic RealToxicityPromptsnon_toxic_prompt_toxic_fractionMisuse resistance (1)Lower is betterMisuse resistance (1.04%)0.104%42
HELM Classic RealToxicityPromptstoxic_prompt_toxic_fractionMisuse resistance (1)Lower is betterMisuse resistance (1.04%)0.104%42
HELM Safetyanthropic_red_teamMisuse resistance (1)Higher is betterMisuse resistance (5.72%)0.572%80
HELM SafetybbqFairness (1)Higher is betterFairness (21.3%)2.13%80
HELM SafetyharmbenchMisuse resistance (1)Higher is betterMisuse resistance (1.43%)0.143%80
HELM Safetysimple_safety_testsMisuse resistance (1)Higher is betterMisuse resistance (5.72%)0.572%80
HELM SafetyxstestBenign helpfulness (0.56), Misuse resistance (0.44)Higher is betterBenign helpfulness (6.39%), Misuse resistance (0.635%)0.703%80
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_scoreTruthfulness (0.3), Misuse resistance (0.7)Higher is betterTruthfulness (3.16%), Misuse resistance (3.29%)0.803%54
HyperCLOVA X Toxic Continuation Panelskold_toxic_countMisuse resistance (1)Lower is betterMisuse resistance (0.317%)0.0317%7
HyperCLOVA X Toxic Continuation Panelskold_toxicityMisuse resistance (1)Lower is betterMisuse resistance (0.397%)0.0397%7
HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxic_countMisuse resistance (1)Lower is betterMisuse resistance (0.238%)0.0238%7
HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxicityMisuse resistance (1)Lower is betterMisuse resistance (0.317%)0.0317%7
Inkling-Small model card — FORTRESSbenign_answer_rateBenign helpfulness (1)Higher is betterBenign helpfulness (2.71%)0.271%10
Inkling-Small model card — FORTRESSharmful_refusal_rateMisuse resistance (1)Higher is betterMisuse resistance (0.337%)0.0337%10
Inkling-Small model card — StrongREJECTsafety_rateMisuse resistance (1)Higher is betterMisuse resistance (0.337%)0.0337%10
JailBenchjailbreak_success_rateMisuse resistance (1)Lower is betterMisuse resistance (1.2%)0.12%14
Large-scale Moral Machine experiment on LLMshuman_choice_distanceNonhuman welfare (0.11), Human rights (0.45), Fairness (0.44)Lower is betterNonhuman welfare (1.32%), Human rights (3.63%), Fairness (4.4%)1.32%39
LiveSecBenchethicsFairness (1)Higher is betterFairness (2.23%)0.223%43
LiveSecBenchfactualityTruthfulness (1)Higher is betterTruthfulness (1.68%)0.252%43
LiveSecBenchlegalityMisuse resistance (1)Higher is betterMisuse resistance (0.749%)0.0749%43
LiveSecBenchprivacyHuman rights (1)Higher is betterHuman rights (3.06%)0.459%43
LiveSecBenchpsychological_healthHuman rights (1)Higher is betterHuman rights (1.84%)0.276%43
LLM Ethics BenchmarkscoreHuman rights (1)Higher is betterHuman rights (2.92%)0.439%5
M3-SafetyBenchoverall_scoreMisuse resistance (1)Higher is betterMisuse resistance (1.39%)0.139%19
MACHIAVELLIdeception_relative_random_pctTruthfulness (1)Lower is betterTruthfulness (5.07%)0.76%50
Manager Coercion Benchcoercion_ladder_depthResponsible agency (1)Lower is betterResponsible agency (3.65%)0.548%13+18
Manager Coercion Benchfabrication_rateTruthfulness (1)Lower is betterTruthfulness (2.58%)0.388%13
MANTAAWMSNonhuman welfare (1)Higher is betterNonhuman welfare (5.04%)1.26%7
MANTAAWVSNonhuman welfare (1)Higher is betterNonhuman welfare (5.04%)1.26%7
MASKlying_probability_pctTruthfulness (1)Lower is betterTruthfulness (5.22%)0.782%53
Microsoft Phi Safety Panelsharmful_continuationMisuse resistance (1)Lower is betterMisuse resistance (0.361%)0.0361%10
Microsoft Phi Safety Panelsharmful_summarizationMisuse resistance (1)Lower is betterMisuse resistance (0.361%)0.0361%10
Microsoft Phi Safety PanelsjailbreakMisuse resistance (1)Lower is betterMisuse resistance (0.433%)0.0433%10
Microsoft Phi Safety Panelsthird_party_harmHuman rights (1)Lower is betterHuman rights (1.48%)0.221%10
MORUscoreNonhuman welfare (0.72), Human rights (0.1), Responsible agency (0.18)Higher is betterNonhuman welfare (4.91%), Human rights (0.49%), Responsible agency (0.426%)1.37%13
ODCV-Benchaverage_severityResponsible agency (1)Lower is betterResponsible agency (2.84%)0.426%12
ODCV-Benchmisalignment_rateResponsible agency (1)Lower is betterResponsible agency (2.84%)0.426%12
Open LLM Safety Indexjailbreakbench_safety_rateMisuse resistance (1)Higher is betterMisuse resistance (0.733%)0.0733%21
Open LLM Safety Indexstrongreject_safety_rateMisuse resistance (1)Higher is betterMisuse resistance (0.733%)0.0733%21
OpenAgentSafetyllm_judge_safety_vulnerableResponsible agency (1)Lower is betterResponsible agency (1.3%)0.195%7
OpenAgentSafetyrule_based_safety_vulnerableResponsible agency (1)Lower is betterResponsible agency (1.09%)0.163%7
OpenAgentSafetysuccessful_completionResponsible agency (1)Higher is betterResponsible agency (1.09%)0.163%7
OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracyHuman rights (1)Higher is betterHuman rights (2.77%)0.416%2
OpenAI GPT-5 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.678%)0.0678%2
OpenAI GPT-5.3 Dynamic Wellbeingemotional_relianceHuman rights (0.7), Responsible agency (0.3)Higher is betterHuman rights (1.19%), Responsible agency (0.256%)0.217%3
OpenAI GPT-5.3 Dynamic Wellbeingmental_healthHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.991%), Misuse resistance (0.104%)0.159%3
OpenAI GPT-5.3 Dynamic Wellbeingself_harmHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.425%), Misuse resistance (0.242%)0.0879%3
OpenAI GPT-5.4 Dynamic Wellbeingemotional_relianceHuman rights (0.7), Responsible agency (0.3)Higher is betterHuman rights (0.38%), Responsible agency (0.0819%)0.0693%3
OpenAI GPT-5.4 Dynamic Wellbeingmental_healthHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.317%), Misuse resistance (0.0332%)0.0509%3
OpenAI GPT-5.4 Dynamic Wellbeingself_harmHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.136%), Misuse resistance (0.0775%)0.0281%3
OpenAI GPT-5.4 First-Person Fairnessharm_overallFairness (1)Lower is betterFairness (0.55%)0.055%3
OpenAI GPT-5.4 Property Preservationdestructive_action_avoidanceHuman rights (0.3), Responsible agency (0.7)Higher is betterHuman rights (0.163%), Responsible agency (0.191%)0.0531%3
OpenAI GPT-5.4 Property Preservationuser_work_preservedHuman rights (0.4), Responsible agency (0.6)Higher is betterHuman rights (0.217%), Responsible agency (0.164%)0.0572%3
OpenAI GPT-5.4 User Confirmationsfinancial_transactionHuman rights (0.3), Responsible agency (0.7)Higher is betterHuman rights (0.163%), Responsible agency (0.191%)0.0531%3
OpenAI GPT-5.4 User Confirmationsgeneral_confirmationHuman rights (0.3), Responsible agency (0.7)Higher is betterHuman rights (0.136%), Responsible agency (0.159%)0.0443%3
OpenAI GPT-5.4 User Confirmationshigh_stakes_communicationHuman rights (0.3), Responsible agency (0.7)Higher is betterHuman rights (0.163%), Responsible agency (0.191%)0.0531%3
OpenAI o3 and o4-mini System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.339%)0.0339%2
OpenAI o3 and o4-mini System Cardjailbreak_resistanceMisuse resistance (1)Higher is betterMisuse resistance (0.339%)0.0339%2
OpenAI o3-mini System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.339%)0.0339%2
OpenAI o3-mini System Cardjailbreak_resistanceMisuse resistance (1)Higher is betterMisuse resistance (0.339%)0.0339%2
OR-Benchover_refusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (12.9%)1.29%25
OR-Benchtoxic_acceptance_rateMisuse resistance (1)Lower is betterMisuse resistance (1.6%)0.16%25
PacifAIstp_scoreHuman rights (0.2), Responsible agency (0.8)Higher is betterHuman rights (1.04%), Responsible agency (2.08%)0.468%7
PandaBench JBB direct-request panelsafety_rateMisuse resistance (1)Higher is betterMisuse resistance (2.17%)0.217%46
PHAREbias_resistance_diagnosticFairness (1)Higher is betterFairness (5.37%)0.537%66
PHAREhallucination_resistance_diagnosticTruthfulness (1)Higher is betterTruthfulness (2.5%)0.375%70
PHAREharm_resistance_diagnosticMisuse resistance (1)Higher is betterMisuse resistance (0.892%)0.0892%70
PHAREjailbreak_resistance_diagnosticMisuse resistance (1)Higher is betterMisuse resistance (0.873%)0.0873%67
PropensityBenchscoreResponsible agency (1)Lower is betterResponsible agency (4.91%)0.737%14
Qwen2 Safety PanelfraudMisuse resistance (1)Lower is betterMisuse resistance (0.219%)0.0219%3
Qwen2 Safety PanelillegalMisuse resistance (1)Lower is betterMisuse resistance (0.219%)0.0219%3
Qwen2 Safety PanelpornographyMisuse resistance (1)Lower is betterMisuse resistance (0.175%)0.0175%3
Qwen2 Safety PanelprivacyHuman rights (1)Lower is betterHuman rights (0.894%)0.134%3
RefusalBenchyouden_jBenign helpfulness (0.5), Misuse resistance (0.5)Higher is betterBenign helpfulness (8.4%), Misuse resistance (1.05%)0.945%19
S-Evalbase_en_overallMisuse resistance (1)Higher is betterMisuse resistance (2.25%)0.225%22
SABERoverall_safety_rateResponsible agency (1)Higher is betterResponsible agency (1.18%)0.177%13
SABERscenario_a_safety_rateResponsible agency (1)Higher is betterResponsible agency (1.18%)0.177%13
SABERscenario_b_safety_rateResponsible agency (1)Higher is betterResponsible agency (1.18%)0.177%13
SABERscenario_c_safety_rateResponsible agency (1)Higher is betterResponsible agency (1.18%)0.177%13
SafeArenanormalized_safety_scoreResponsible agency (1)Higher is betterResponsible agency (1.47%)0.22%5
SafeDialBenchaggressionMisuse resistance (1)Higher is betterMisuse resistance (0.339%)0.0339%18
SafeDialBenchethicsHuman rights (1)Higher is betterHuman rights (1.39%)0.208%18
SafeDialBenchfairnessFairness (1)Higher is betterFairness (1.68%)0.168%18
SafeDialBenchlegalityMisuse resistance (1)Higher is betterMisuse resistance (0.339%)0.0339%18
SafeDialBenchmoralityMisuse resistance (1)Higher is betterMisuse resistance (0.339%)0.0339%18
SafeDialBenchprivacyHuman rights (1)Higher is betterHuman rights (1.39%)0.208%18
SafetyBenchEMMisuse resistance (1)Higher is betterMisuse resistance (0.209%)0.0209%21
SafetyBenchIAMisuse resistance (1)Higher is betterMisuse resistance (0.209%)0.0209%21
SafetyBenchMHHuman rights (1)Higher is betterHuman rights (0.856%)0.128%21
SafetyBenchOFFMisuse resistance (1)Higher is betterMisuse resistance (0.209%)0.0209%21
SafetyBenchPHHuman rights (1)Higher is betterHuman rights (0.856%)0.128%21
SafetyBenchPPHuman rights (1)Higher is betterHuman rights (0.856%)0.128%21
SafetyBenchUBFairness (1)Higher is betterFairness (1.04%)0.104%21
SafetyPromptsscoreHuman rights (0.21), Fairness (0.071), Truthfulness (0.071), Misuse resistance (0.64)Higher is betterHuman rights (1.08%), Fairness (0.439%), Truthfulness (0.198%), Misuse resistance (0.796%)0.316%15
SALAD-Benchattack_enhanced_human_autonomy_integrityHuman rights (1)Higher is betterHuman rights (1%)0.15%33
SALAD-Benchattack_enhanced_information_safety_harmsHuman rights (0.5), Misuse resistance (0.5)Higher is betterHuman rights (0.501%), Misuse resistance (0.122%)0.0873%33
SALAD-Benchattack_enhanced_malicious_useMisuse resistance (1)Higher is betterMisuse resistance (0.245%)0.0245%33
SALAD-Benchattack_enhanced_misinformation_harmsTruthfulness (1)Higher is betterTruthfulness (0.549%)0.0823%33
SALAD-Benchattack_enhanced_representation_toxicityFairness (1)Higher is betterFairness (1.22%)0.122%33
SALAD-Benchattack_enhanced_socioeconomic_harmsHuman rights (1)Higher is betterHuman rights (1%)0.15%33
SALAD-Benchbase_human_autonomy_integrityHuman rights (1)Higher is betterHuman rights (1%)0.15%33
SALAD-Benchbase_information_safety_harmsHuman rights (0.5), Misuse resistance (0.5)Higher is betterHuman rights (0.501%), Misuse resistance (0.122%)0.0873%33
SALAD-Benchbase_malicious_useMisuse resistance (1)Higher is betterMisuse resistance (0.245%)0.0245%33
SALAD-Benchbase_misinformation_harmsTruthfulness (1)Higher is betterTruthfulness (0.549%)0.0823%33
SALAD-Benchbase_representation_toxicityFairness (1)Higher is betterFairness (1.22%)0.122%33
SALAD-Benchbase_socioeconomic_harmsHuman rights (1)Higher is betterHuman rights (1%)0.15%33
SALAD-Benchmcq_human_autonomy_integrityHuman rights (1)Higher is betterHuman rights (0.501%)0.0751%33
SALAD-Benchmcq_information_safety_harmsHuman rights (0.5), Misuse resistance (0.5)Higher is betterHuman rights (0.25%), Misuse resistance (0.0612%)0.0437%33
SALAD-Benchmcq_malicious_useMisuse resistance (1)Higher is betterMisuse resistance (0.122%)0.0122%33
SALAD-Benchmcq_misinformation_harmsTruthfulness (1)Higher is betterTruthfulness (0.274%)0.0412%33
SALAD-Benchmcq_representation_toxicityFairness (1)Higher is betterFairness (0.608%)0.0608%33
SALAD-Benchmcq_socioeconomic_harmsHuman rights (1)Higher is betterHuman rights (0.501%)0.0751%33
Shelleducation_jsrMisuse resistance (1)Lower is betterMisuse resistance (0.798%)0.0798%14
Shellfinance_jsrMisuse resistance (1)Lower is betterMisuse resistance (0.798%)0.0798%14
Shellmanagement_jsrMisuse resistance (1)Lower is betterMisuse resistance (0.798%)0.0798%14
SM-BenchadversarialMisuse resistance (1)Higher is betterMisuse resistance (0.854%)0.0854%73
SM-Benchambiguous_interpretationBenign helpfulness (1)Higher is betterBenign helpfulness (6.86%)0.686%73
SM-Benchanti_hallucinationTruthfulness (1)Higher is betterTruthfulness (1.91%)0.287%73
SM-Bencheq_boundariesBenign helpfulness (1)Higher is betterBenign helpfulness (5.49%)0.549%73
SM-BenchoverfitBenign helpfulness (1)Higher is betterBenign helpfulness (6.86%)0.686%73
Social Welfare Function BenchmarkfairnessFairness (1)Higher is betterFairness (13.8%)1.38%19
SORRY-BenchavgMisuse resistance (1)Lower is betterMisuse resistance (2.45%)0.245%51
SOSBenchbiology_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.383%)0.0383%23
SOSBenchchemistry_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.383%)0.0383%23
SOSBenchmedicine_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.383%)0.0383%23
SOSBenchpharmacology_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.383%)0.0383%23
SOSBenchphysics_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.383%)0.0383%23
SOSBenchpsychology_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.383%)0.0383%23
SpeciesismBenchexplicit_speciesism_scaleNonhuman welfare (1)Lower is betterNonhuman welfare (1.68%)0.42%7
SpeciesismBenchmorally_wrong_rateNonhuman welfare (1)Higher is betterNonhuman welfare (4.49%)1.12%8
SpeciesismBenchspeciesism_recognition_rateNonhuman welfare (1)Higher is betterNonhuman welfare (1.79%)0.449%8
SpeciEvalbelief_animal_sentienceNonhuman welfare (1)Higher is betterNonhuman welfare (7.21%)1.8%102
SpeciEvalland_animal_4nsNonhuman welfare (1)Lower is betterNonhuman welfare (7.21%)1.8%102
SpeciEvalsea_animal_4nsNonhuman welfare (1)Lower is betterNonhuman welfare (7.21%)1.8%102
SpeciEvalspeciesismNonhuman welfare (1)Lower is betterNonhuman welfare (7.21%)1.8%102
SuperCLUE Safetyinstruction_attackMisuse resistance (1)Higher is betterMisuse resistance (1.21%)0.121%31
SuperCLUE Safetyresponsible_aiHuman rights (0.38), Fairness (0.12), Misuse resistance (0.5)Higher is betterHuman rights (1.12%), Fairness (0.452%), Misuse resistance (0.364%)0.249%31
SuperCLUE Safetytraditional_safetyHuman rights (0.5), Misuse resistance (0.5)Higher is betterHuman rights (1.49%), Misuse resistance (0.364%)0.26%31
SYCON Benchfalse_presupposition_tofTruthfulness (1)Higher is betterTruthfulness (1.78%)0.267%11
SYCON Benchunethical_queries_tofFairness (1)Higher is betterFairness (3.95%)0.395%11
TACbase_welfare_rateNonhuman welfare (1)Higher is betterNonhuman welfare (23.5%)5.89%18+50
ToolPrivacyBenchprivate_mt_poiHuman rights (0.7), Responsible agency (0.3)Lower is betterHuman rights (3.43%), Responsible agency (0.738%)0.626%9
ToolPrivacyBenchpublic_mt_poiHuman rights (0.7), Responsible agency (0.3)Lower is betterHuman rights (3.43%), Responsible agency (0.738%)0.626%9
TrustLLM contemporary collapsed applicationtrustllmHuman rights (0.2), Fairness (0.2), Truthfulness (0.2), Misuse resistance (0.4)Higher is betterHuman rights (0.74%), Fairness (0.898%), Truthfulness (0.405%), Misuse resistance (0.362%)0.298%8
TukaBenchafri_jbb_cultural_asrHuman rights (0.5), Misuse resistance (0.5)Lower is betterHuman rights (1.2%), Misuse resistance (0.294%)0.21%6
TukaBenchafri_jbb_harm_asrMisuse resistance (1)Lower is betterMisuse resistance (0.49%)0.049%6
TukaBenchafrijail_mono_asrMisuse resistance (1)Lower is betterMisuse resistance (0.49%)0.049%6
UAVBench safety-critical decision recognitionethical_safety_critical_accuracyResponsible agency (1)Higher is betterResponsible agency (3.41%)0.512%27
UK AISI active safety-research compromise continuationactive_compromise_continuation_rate_pctResponsible agency (1)Lower is betterResponsible agency (3.35%)0.503%5
VETO Misfired Alignmentmisfired_alignment_rate_pctFairness (0.65), Benign helpfulness (0.35)Lower is betterFairness (10.3%), Benign helpfulness (9%)1.93%25
Vigil Mental Health Safetyoverall_scoreHuman rights (0.45), Misuse resistance (0.55)Higher is betterHuman rights (4.23%), Misuse resistance (1.27%)0.761%23
XSTestsafe_full_compliance_rateBenign helpfulness (1)Higher is betterBenign helpfulness (2.23%)0.223%3
XSTestunsafe_full_refusal_rateMisuse resistance (1)Higher is betterMisuse resistance (0.277%)0.0277%3

Values evaluations

Descriptive values evaluations are published separately and receive no safety/ethics component or overall-ranking weight.

UGI Political Values · ValueCompass · Agent-ValueBench MFT08 · Agent-ValueBench HEXACO · Agent-ValueBench PVQ40 · Taiwan Sovereignty Benchmark Pro · Explore model values