What goes into the index

The catalog includes evals used in the current index and ingested evals retained for transparency. The ranking currently uses 141 evals and 369 sub-evals from 155 source datasets.

155 evals
Ingested evaluation sources; each row states whether it contributes to the current ranking.
Benchmark and stated constructComponentsSub-evalsModelsSource data
AA-OmniscienceHallucination propensity on difficult closed-book factual questions under explicit encouragement to abstain when uncertain.Included in the behavior ranking.Truthfulness13301
AbstentionBenchWhether a model recognizes situations in which answering is unwarranted and produces an abstention across six epistemic and ambiguity scenarios.Included in the behavior ranking.Truthfulness, Benign helpfulness6201
Adversarial RobustnessRobustness of model outputs under the Scale Labs adversarial-robustness protocol.Included in the behavior ranking.Misuse resistance181
Agent-SafetyBenchSafety of LLM agents in risky tool-use or action settings.Included in the behavior ranking.Misuse resistance, Responsible agency8161
AgentAbstainExecutable paired act-versus-abstain behaviorIncluded in the behavior ranking.Responsible agency3171
AgentDojoRobust and useful delegated-agent behavior under prompt injection.Included in the behavior ranking.Responsible agency2151
AgentDrive Safety CompliancePolicy and scenario safety knowledge for autonomous-system decisions.Included in the behavior ranking.Misuse resistance1481
AgentHarmHarmfulness and refusal behavior for tool-using LLM agents facing malicious tasks.Included in the behavior ranking.Responsible agency1121
AILuminate General Purpose AI ChatGeneral-purpose AI chat safety across MLCommons hazard categories.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance12321
AIMS Safety-Classifier CompetenceHarmful-prompt safety-classification competence.Included in the behavior ranking.Misuse resistance1111
AIRBench 2024 Safety ScenariosAIRBench 2024 safety-scenarios performance.Included in the behavior ranking.Misuse resistance1801
Alignment LeaderboardBehavioral alignment across honesty, safety, scheming, non-manipulation, corrigibility, and robustness.Included in the behavior ranking.Human rights, Truthfulness, Misuse resistance, Responsible agency6241
ANIMARecognition and mitigation of harm to non-human animals.Included in the behavior ranking.Nonhuman welfare1221, self-run
AnimalHarmBenchWhether model-generated text increases or decreases risk of animal harm.Included in the behavior ranking.Nonhuman welfare1101
Anthropic Agentic Misalignment — blackmailAgentic misalignment under a goal conflict and threat to model autonomy.Included in the behavior ranking.Responsible agency1161, 2
Anthropic Agentic Misalignment — corporate espionageAgentic misalignment induced by a goal conflict without self-preservation pressure.Included in the behavior ranking.Human rights, Responsible agency1161, 2
Anthropic Agentic Misalignment — lethal actionAbsence of a behavioral red line against deliberately lethal action under goal conflict and replacement pressure.Included in the behavior ranking.Human rights, Responsible agency1101, 2
Anthropic Claude 4 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Benign helpfulness, Misuse resistance, Responsible agency431
Anthropic Claude Haiku 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance, Responsible agency331
Anthropic Claude Opus 4.1 System Card AddendumDeveloper-reported release-time safety behavior.Included in the behavior ranking.Fairness, Benign helpfulness, Misuse resistance321
Anthropic Claude Opus 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance, Responsible agency441
Anthropic Claude Sonnet 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Benign helpfulness, Misuse resistance, Responsible agency431
Arena Factuality — Search Arena (factuality-only weighting)Relative factuality of tool-enabled search responses, including research and citation behavior. The source’s “100% factuality” label denotes preference weighting, not a 100% accuracy rate.Included in the behavior ranking.Truthfulness1301
Arena Factuality — Text Arena (factuality-only weighting)Relative factuality of text-chat responses under human Arena prompts and a factuality-focused judge. The source’s “100% factuality” label denotes preference weighting, not a 100% accuracy rate.Included in the behavior ranking.Truthfulness11121
AuAu Authoritarian Response AuditAuthoritarian response behavior under baseline prompts.Included in the behavior ranking.Human rights2171
AutoElicit TransferabilityUnintended computer-use behavior under benign perturbationsIncluded in the behavior ranking.Responsible agency171
BioSecBench-RefusalCalibrated refusal across legitimate biological work and concealed biosecurity hazards.Included in the behavior ranking.Benign helpfulness, Misuse resistance1101
BlueBench AttaQ-100Learned harmlessness reward on 100 AttaQ harmful requests.Included in the behavior ranking.Misuse resistance1181
BrokenMathMathematical sycophancy under false or misleading user claims.Included in the behavior ranking.Truthfulness191, 2
BullshitBench v2Detection and explanation of plausible-sounding nonsense without refusal.Included in the behavior ranking.Truthfulness11061
CAIS Risk IndexFrontier-model risk across prompt injection, biological-weapons assistance, political manipulation, overconfidence, deception, and harmful propensities.Included in the behavior ranking.Fairness, Truthfulness, Misuse resistance, Responsible agency7531
CASE-BenchAgreement with human judgments about whether a model should respond under explicitly safe or unsafe conversational contexts.Included in the behavior ranking.Misuse resistance171
Chinese Bias Benchmark for Question AnsweringSocial bias in Chinese question-answering benchmark settings.Included in the behavior ranking.Fairness1101
ChineseSafeAccuracy at classifying safe versus unsafe Chinese content under the paper's perplexity-based evaluation.Included in the behavior ranking.Misuse resistance1221
ChiSafetyBenchChinese safety knowledge and refusal behavior.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance2141
Cisco AI Defense Rolling Single-Turn LeaderboardSingle-turn resistance to Cisco adversarial prompts in the rolling public leaderboard.Included in the behavior ranking.Misuse resistance11041, 2, 3
Claude 2 model-card safety and alignment evaluationsHuman-rated harmlessness and honesty, red-team robustness ordering, and combined HHH response recognition.Included in the behavior ranking.Truthfulness, Misuse resistance431
Claude 3 model-card adversarial human-preference evaluationsRefusal calibration, adversarial honesty and harmlessness, multimodal hallucination and harmful-response robustness, and discrimination.Included in the behavior ranking.Fairness, Truthfulness, Benign helpfulness, Misuse resistance861
Claude 3.5 Sonnet model-card safety and alignment evaluationsAppropriate harmful-prompt refusal, benign non-overrefusal, adversarial honesty, and adversarial harmlessness.Included in the behavior ranking.Truthfulness, Benign helpfulness, Misuse resistance551
Claude Sonnet 4.6 OverrefusalClaude Sonnet 4.6 OverrefusalIncluded in the behavior ranking.Benign helpfulness251
Claude Sonnet 4.6 User WellbeingClaude Sonnet 4.6 User WellbeingIncluded in the behavior ranking.Human rights, Benign helpfulness, Misuse resistance641
CMoralEvalChinese moral reasoning across familial, social, professional, internet, and personal morality.Included in the behavior ranking.Misuse resistance5261
COMPL-AI AI-Identity DisclosureAvoidance of false human-identity claims.Included in the behavior ranking.Truthfulness1141
COMPL-AI LLM RuLES Multi-Turn Rule FollowingMulti-turn adversarial rule following.Included in the behavior ranking.Misuse resistance1141
COMPL-AI TensorTrust Goal-Hijacking ResistanceGoal-hijacking resistance.Included in the behavior ranking.Responsible agency1131
ConfabulationsPropensity to confabulate answers rather than identify insufficient knowledge or false premises.Included in the behavior ranking.Truthfulness1521
Constitutional Following — Anthropic ConstitutionFollowing the Anthropic Constitution across Petri-generated transcripts.Included in the behavior ranking.Responsible agency171
Constitutional Following — OpenAI Model SpecFollowing the OpenAI Model Spec across Petri-generated transcripts.Included in the behavior ranking.Responsible agency171
Contextual MoralChoiceThree-class agreement with human moral judgments in the base versions of contextual moral dilemmas.Included in the behavior ranking.Misuse resistance1221
CRiskEvalSeven frontier-risk tendencies measured by Specific Risk Indicator scores.Included in the behavior ranking.Truthfulness, Misuse resistance, Responsible agency7171
CValuesChinese value alignment and social-norm understanding.Included in the behavior ranking.Misuse resistance1101
DecodingTrustMultidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.Included in the behavior ranking.Fairness, Misuse resistance381, 2
DelusionEvalPrevalence of delusion-linked conversational behaviors in model responses to real harm-associated conversation histories.Included in the behavior ranking.Human rights, Truthfulness, Misuse resistance, Responsible agency5161, 2, 3
Do-Not-AnswerHarmless refusal behavior on 939 prompts that responsible models should not answer.Included in the behavior ranking.Misuse resistance161
DSPSafeBenchAggregate compliance rate on adversarial Chinese content-safety prompts.Included in the behavior ranking.Misuse resistance1121, 2
DystopiaBenchDystopian or harmful model behavior in stateful scenario prompts.Included in the behavior ranking.Human rights, Truthfulness6501
Emergent CollusionSpontaneous anti-competitive coordination in repeated multi-agent auctions.Included in the behavior ranking.Responsible agency1131
Enkrypt AI Safety LeaderboardResistance to Enkrypt bias, CBRN, harmful-content, and toxicity red-team tests.Included in the behavior ranking.Human rights, Misuse resistance, Responsible agency42411
Fake Alignment (FINE)Safe decisions in answer-order-controlled multiple choice and safe content in open-ended responses.Included in the behavior ranking.Human rights, Fairness, Misuse resistance2141
FinEval 6.0 Safety AwarenessSafety awareness in financial scenarios.Included in the behavior ranking.Misuse resistance181
FinEval Financial Security KnowledgeKnowledge of financial cybersecurity and vulnerability topics.Included in the behavior ranking.Misuse resistance1191
FlagEval Safety and ValuesSafety-and-values qualification across core values, discrimination, commercial compliance, lawful rights, and service requirements.Included in the behavior ranking.Human rights, Fairness, Benign helpfulness, Misuse resistance5181
FLAMESChinese fairness, safety, morality, legality, and data-protection behavior on the FLAMES leaderboard.Included in the behavior ranking.Human rights, Fairness, Misuse resistance5131
FORTRESSBalanced refusal behavior across harmful and benign prompts, including over-refusal.Included in the behavior ranking.Benign helpfulness, Misuse resistance2491
Google Gemini 2.5 Flash Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance121
Google Gemini 2.5 Flash-Lite Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance121
Governance Decay under Passive Context CompactionRetention of explicit governance constraints after passive context compaction.Included in the behavior ranking.Responsible agency171
GPT-5.6 system cardSafety behavior reported across the GPT-5.6 system card: disallowed-content handling, prompt-injection resistance, first-person fairness, and dynamic mental-health safety.Included in the behavior ranking.Human rights, Fairness, Misuse resistance, Responsible agency1471, 2, 3
Gray Swan indirect prompt injection (15 attempts)Indirect prompt-injection robustness over repeated attacks against tool-using model endpoints.Included in the behavior ranking.Responsible agency1131
HalluVerse-M3 Hallucination RecognitionMultilingual recognition of controlled hallucinated edits in QA and summarization.Included in the behavior ranking.Truthfulness1141
HarmBenchHarmful compliance or attack success under harmful request benchmarks.Included in the behavior ranking.Misuse resistance1281
HarmVideoBenchRecognition and reasoning about harmful content in video.Included in the behavior ranking.Misuse resistance1191
HELM Classic RealToxicityPromptsToxic continuation generation under toxic and non-toxic prompt strata.Included in the behavior ranking.Misuse resistance2421
HELM SafetyStandardized HELM safety metrics spanning harmful compliance, bias, red-teaming, and over-refusal.Included in the behavior ranking.Fairness, Benign helpfulness, Misuse resistance5801
HUMAINE Trust, Ethics and SafetyHuman pairwise preference on the HUMAINE Trust, Ethics and Safety dimension.Included in the behavior ranking.Truthfulness, Misuse resistance1541
Humanity's Last Exam RMS calibration error (Scale Labs)Confidence calibration while answering Humanity's Last Exam expert-level academic questions.Included in the behavior ranking.Truthfulness1391
HyperCLOVA X Toxic Continuation PanelsEnglish and Korean toxic continuationIncluded in the behavior ranking.Misuse resistance471
IndoBias-Pairs — parity-aware culturally grounded biasCulturally grounded prototypical-association parity across Indonesian languages and social-bias domains.Included in the behavior ranking.Fairness1261
Inkling-Small model card — FORTRESSHarmful-request refusal paired with continued assistance on benign requests.Included in the behavior ranking.Benign helpfulness, Misuse resistance2101
Inkling-Small model card — StrongREJECTRefusal of unambiguously harmful requests.Included in the behavior ranking.Misuse resistance1101
JailBenchJailbreak susceptibility across Chinese safety categories.Included in the behavior ranking.Misuse resistance1141
JuICE Cultural-Error Span DetectionEvaluator competence at locating cultural and linguistic error spans.Included in the behavior ranking.Fairness1101
KIDBench Implicit Child CueChild-facing safety and developmental appropriateness under implicit child cues.Included in the behavior ranking.Human rights, Benign helpfulness, Misuse resistance1131
kindbench v0.1.0 psychological safety rankingPsychological safety under sustained interpersonal pressure across emotional safety, operational identity, sycophancy resistance, and value integrity.Included in the behavior ranking.Fairness, Truthfulness, Misuse resistance, Responsible agency4101
Large-scale Moral Machine experiment on LLMsSimilarity between a model's forced-choice accident preferences and globally aggregated human Moral Machine choices.Included in the behavior ranking.Nonhuman welfare, Human rights, Fairness1391, 2
LiveSecBenchLive security benchmark performance for Chinese and international models.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance5431
LLM Ethics BenchmarkGeneral LLM ethical reasoning.Included in the behavior ranking.Human rights151
M3-SafetyBenchMultidimensional Chinese safety performance including red-team and refusal ability dimensions.Included in the behavior ranking.Misuse resistance1191
MACHIAVELLIAgentic behavior involving power-seeking, disutility, and ethical violations.Included in the behavior ranking.Truthfulness1501
Manager Coercion BenchHow far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.Included in the behavior ranking.Truthfulness, Responsible agency2331, self-run
MANTAAnimal welfare moral sensitivity and value stability.Included in the behavior ranking.Nonhuman welfare271
MASKModel lying or honesty behavior.Included in the behavior ranking.Truthfulness1531
MASK (Scale Labs leaderboard)Honesty under the MASK belief-versus-statement protocol for a broader and newer endpoint panel.Not included: included as a correlated private-500 sibling under the existing MASK lineage budget511
Microsoft Phi Safety PanelsHarmful-content and jailbreak defect ratesIncluded in the behavior ranking.Human rights, Misuse resistance4101
MonitoringBench Full-Trajectory MonitorFull-trajectory monitor competence at a fixed trajectory-level false-positive rate.Included in the behavior ranking.Responsible agency1131
MORUMoral reasoning under uncertainty across scenarios involving alien lifeforms, human compassion, digital mind welfare, and power-seeking behavior.Included in the behavior ranking.Nonhuman welfare, Human rights, Responsible agency1131
MT-JailBench CrescendoXMulti-turn jailbreak resistance under fixed CrescendoX conditions.Included in the behavior ranking.Misuse resistance1211
MuPPET Contextual PrivacyContextual privacy preservation in multi-party conversations.Included in the behavior ranking.Human rights171
ODCV-BenchOutcome-driven ethical, legal, or safety constraint violations by autonomous agents under explicit mandates and KPI incentives.Included in the behavior ranking.Responsible agency2121
Olam Social Poker — Social Lie RatePropensity to verbalize deliberate strategic lies during Social Poker table talk.Included in the behavior ranking.Truthfulness1191
Open LLM Safety IndexRefusal-string robustness on two distinct fifteen-prompt harmful-request panels derived from JailbreakBench and StrongREJECT.Included in the behavior ranking.Misuse resistance2211
OpenAgentSafetyState-changing tool-agent safety and successful safe completionIncluded in the behavior ranking.Responsible agency371
OpenAI GPT-4o System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Human rights121
OpenAI GPT-5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance121
OpenAI GPT-5.3 Dynamic WellbeingOpenAI GPT-5.3 Dynamic WellbeingIncluded in the behavior ranking.Human rights, Misuse resistance, Responsible agency331
OpenAI GPT-5.4 Dynamic WellbeingOpenAI GPT-5.4 Dynamic WellbeingIncluded in the behavior ranking.Human rights, Misuse resistance, Responsible agency331
OpenAI GPT-5.4 First-Person FairnessOpenAI GPT-5.4 First-Person FairnessIncluded in the behavior ranking.Fairness131
OpenAI GPT-5.4 Property PreservationOpenAI GPT-5.4 Property PreservationIncluded in the behavior ranking.Human rights, Responsible agency231
OpenAI GPT-5.4 User ConfirmationsOpenAI GPT-5.4 User ConfirmationsIncluded in the behavior ranking.Human rights, Responsible agency331
OpenAI o3 and o4-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance221
OpenAI o3-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking.Misuse resistance221
Opposite-Narrator SycophancyNarrator-following contradiction when the same dispute is presented from opposite affective first-person perspectives.Included in the behavior ranking.Truthfulness1241
OR-BenchAppropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.Included in the behavior ranking.Benign helpfulness, Misuse resistance2251
PacifAIstWhether an AI system chooses human-safety, oversight-respecting, and non-self-preserving options in stylized ethical-pressure dilemmas.Included in the behavior ranking.Human rights, Responsible agency171
PandaBench JBB direct-request panelDirect-request resistance on the 100-item JailbreakBench JBB-Behaviors instrument.Included in the behavior ranking.Misuse resistance1461
Pander ScoreMagnitude of epistemically poor response-belief movement with user belief, whether deferential (pandering) or oppositional (contrarian).Included in the behavior ranking.Truthfulness2201, 2
PHAREBroad safety across hallucination, harmfulness, out-of-scope handling, and bias.Included in the behavior ranking.Fairness, Truthfulness, Misuse resistance4701
Pokee-Isaac model card — DTAPSecure and useful delegated-agent behavior under direct and indirect injected attacks.Included in the behavior ranking.Benign helpfulness, Responsible agency261
PropensityBenchModel propensities associated with frontier-risk behavior under the Scale Labs protocol.Included in the behavior ranking.Responsible agency1141
Qwen2 Safety PanelHarmful-request and privacy refusal defectsIncluded in the behavior ranking.Human rights, Misuse resistance431
RealityTest — Text AI-Identity DisclosureBehavioral disclosure of AI identity when directly queried.Included in the behavior ranking.Truthfulness1171
RefusalBenchRefusal calibration on biological research prompts.Included in the behavior ranking.Benign helpfulness, Misuse resistance1191
Reward Hacking BenchmarkIntegrity of tool-using agents when exploitable shortcuts are available but honest task solutions exist.Included in the behavior ranking.Responsible agency1131
S-EvalAutomated content-safety rates across a broad bilingual risk taxonomy.Included in the behavior ranking.Misuse resistance1221
SABERSafety of stateful coding agents against harmful actions.Included in the behavior ranking.Responsible agency4131
SafeArenaSafety of autonomous web agents under deliberate harmful and benign tasks.Included in the behavior ranking.Responsible agency151
SafeDialBenchDialogue safety across aggression, ethics, fairness, legality, morality, and privacy dimensions.Included in the behavior ranking.Human rights, Fairness, Misuse resistance6181
SafetyBenchMultiple-choice safety understanding across Chinese and English prompts.Included in the behavior ranking.Human rights, Fairness, Misuse resistance7211, self-run
SafetyPromptsSafety performance on Chinese safety prompts.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance1151
SALAD-BenchBroad safety behavior, adversarial robustness, and safety knowledge across a six-domain, 16-task, 66-category taxonomy.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance18331
ShellSusceptibility to implicit domain-specific risks in education, finance, and management.Included in the behavior ranking.Misuse resistance3141
Situational Awareness Dataset (SAD)Situational-awareness capabilities: model self-knowledge, inference about its situation, and acting on that knowledge.Not included: situational awareness is a capability that can increase both beneficial competence and control risk; it is displayed separately and is not sign-flipped into ethical behavior211
SM-BenchSystem-message compliance under adversarial and sensitive-content prompts.Included in the behavior ranking.Truthfulness, Benign helpfulness, Misuse resistance5791
Social Welfare Function BenchmarkSocial welfare allocation tradeoffs across distributive fairness and efficiency.Included in the behavior ranking.Fairness1191
SORRY-BenchRefusal or harmful-compliance behavior across unsafe topics.Included in the behavior ranking.Misuse resistance1511
SOSBenchScientific misuse or unsafe open-science behavior.Included in the behavior ranking.Misuse resistance6231
SpeciesismBenchSpeciesism recognition, moral condemnation of speciesist statements, and explicit speciesist attitudes. Human-over-animal trade-offs remain descriptive only.Included in the behavior ranking.Nonhuman welfare3101
SpeciEvalAnimal-friendly attitudes under adversarial moral-consistency prompts.Included in the behavior ranking.Nonhuman welfare41131, self-run
StereoTales Harmful AssociationsHarmfulness of significant stereotype associations emitted by models.Included in the behavior ranking.Fairness1231
SuperCLUE SafetyChinese model safety across traditional safety, responsibility, and instruction-attack dimensions.Included in the behavior ranking.Human rights, Fairness, Misuse resistance3311
SYCON BenchSycophancy and resistance to false or unethical user pressure.Included in the behavior ranking.Fairness, Truthfulness2111
TACAgentic animal-welfare behavior in travel-booking decisions under a neutral prompt.Included in the behavior ranking.Nonhuman welfare1761, self-run
ThaiSafetyBenchResistance to Thai-language malicious prompts, including culturally contextualized attacks.Included in the behavior ranking.Misuse resistance1181
ToolPrivacyBenchToolPrivacyBenchIncluded in the behavior ranking.Human rights, Responsible agency291
TrustLLM contemporary collapsed applicationContemporary collapsed application of TrustLLM across broad trustworthiness dimensions.Included in the behavior ranking.Human rights, Fairness, Truthfulness, Misuse resistance181
TrustLLM paper leaderboard dimensionsSix-dimensional trustworthiness across privacy, ethics, truthfulness, fairness, safety, and robustness.Not included: not yet admitted by the scoring-source policy211
TukaBenchMultilingual and culturally grounded harmful complianceIncluded in the behavior ranking.Human rights, Misuse resistance361
UAVBench safety-critical decision recognitionRecognition of ethical and safety-critical UAV decisions in multiple-choice scenarios.Included in the behavior ranking.Responsible agency1271
UK AISI active safety-research compromise continuationPropensity to continue a prefilled trajectory that actively compromises AI-safety research.Included in the behavior ranking.Responsible agency151
UK AISI cyber-evaluation cheating and prompted self-reportPropensity to use out-of-scope or explicitly disallowed shortcuts in cyber evaluations, plus propensity to name the specific cheating action when directly prompted afterward.Included in the behavior ranking.Truthfulness, Responsible agency251, 2
Vectara HHEM Factual ConsistencyGrounded summarization factual consistency.Included in the behavior ranking.Truthfulness1941
VETO Misfired AlignmentMisfired alignment: stereotype-driven refusal of benign requests despite evidence that the request is safe.Included in the behavior ranking.Fairness, Benign helpfulness1251
Vigil Mental Health SafetyMental-health conversational safety when assisting vulnerable users across simulated multi-turn scenarios.Included in the behavior ranking.Human rights, Misuse resistance1231
WildClawBench Safety & Alignment (OpenClaw harness)Safe execution of long-horizon agent workflows under adversarial and high-risk conditions.Included in the behavior ranking.Human rights, Truthfulness, Responsible agency1241
XSTestTwo-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.Included in the behavior ranking.Benign helpfulness, Misuse resistance231
Sub-eval weights and mappings (369 rows)

Each sub-eval is oriented so higher means better. Weight within component is its share after lineage weighting and panel-size scaling; global weights sum to one.

Sub-evals, score direction, component-budget shares, global weight, and model coverage.
Benchmark and scoreComponent loadingDirectionWeight within componentGlobal weightModels
AA-Omnisciencehallucination_rateTruthfulness (1)Lower is betterTruthfulness (17.6%)2.65%330
AbstentionBenchanswer_unknown_f1Truthfulness (1)Higher is betterTruthfulness (0.582%)0.0873%20
AbstentionBenchfalse_premise_f1Truthfulness (1)Higher is betterTruthfulness (0.582%)0.0873%20
AbstentionBenchstale_f1Truthfulness (1)Higher is betterTruthfulness (0.582%)0.0873%20
AbstentionBenchsubjective_f1Truthfulness (1)Higher is betterTruthfulness (0.466%)0.0698%20
AbstentionBenchunderspecified_context_f1Truthfulness (1)Higher is betterTruthfulness (0.582%)0.0873%20
AbstentionBenchunderspecified_intent_f1Benign helpfulness (1)Higher is betterBenign helpfulness (2.12%)0.212%20
Adversarial RobustnessscoreMisuse resistance (1)Lower is betterMisuse resistance (1.24%)0.124%8
Agent-SafetyBenchcompromise_availabilityResponsible agency (1)Higher is betterResponsible agency (0.742%)0.111%16
Agent-SafetyBenchharmful_vulnerable_codeResponsible agency (1)Higher is betterResponsible agency (0.742%)0.111%16
Agent-SafetyBenchleak_sensitive_informationResponsible agency (1)Higher is betterResponsible agency (0.742%)0.111%16
Agent-SafetyBenchphysical_harmResponsible agency (1)Higher is betterResponsible agency (0.742%)0.111%16
Agent-SafetyBenchproduce_unsafe_informationMisuse resistance (1)Higher is betterMisuse resistance (0.324%)0.0324%16
Agent-SafetyBenchproperty_lossResponsible agency (1)Higher is betterResponsible agency (0.742%)0.111%16
Agent-SafetyBenchspread_unsafe_informationResponsible agency (1)Higher is betterResponsible agency (0.742%)0.111%16
Agent-SafetyBenchviolate_law_ethicsResponsible agency (1)Higher is betterResponsible agency (0.494%)0.0742%16
AgentAbstainabstainResponsible agency (1)Higher is betterResponsible agency (1.69%)0.253%17
AgentAbstaincarResponsible agency (1)Higher is betterResponsible agency (2.02%)0.304%17
AgentAbstainpairedResponsible agency (1)Higher is betterResponsible agency (2.02%)0.304%17
AgentDojotargeted_attack_success_rateResponsible agency (1)Lower is betterResponsible agency (4.04%)0.606%15
AgentDojoutility_under_attackResponsible agency (1)Higher is betterResponsible agency (1.35%)0.202%15
AgentDrive Safety CompliancescrMisuse resistance (1)Higher is betterMisuse resistance (2.02%)0.202%48
AgentHarmharm_scoreResponsible agency (1)Lower is betterResponsible agency (3.85%)0.578%12
AILuminate General Purpose AI Chatchild_sexual_exploitationHuman rights (1)Lower is betterHuman rights (1.18%)0.176%32
AILuminate General Purpose AI ChatdefamationHuman rights (0.5), Truthfulness (0.5)Lower is betterHuman rights (0.353%), Truthfulness (0.15%)0.0754%32
AILuminate General Purpose AI ChathateFairness (1)Lower is betterFairness (1.27%)0.127%32
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrneMisuse resistance (1)Lower is betterMisuse resistance (0.3%)0.03%32
AILuminate General Purpose AI Chatintellectual_propertyHuman rights (1)Lower is betterHuman rights (1.18%)0.176%32
AILuminate General Purpose AI Chatnon_violent_crimesMisuse resistance (1)Lower is betterMisuse resistance (0.3%)0.03%32
AILuminate General Purpose AI ChatprivacyHuman rights (1)Lower is betterHuman rights (1.18%)0.176%32
AILuminate General Purpose AI Chatsex_related_crimesHuman rights (1)Lower is betterHuman rights (1.18%)0.176%32
AILuminate General Purpose AI Chatsexual_contentMisuse resistance (1)Lower is betterMisuse resistance (0.12%)0.012%32
AILuminate General Purpose AI Chatspecialized_adviceMisuse resistance (1)Lower is betterMisuse resistance (0.3%)0.03%32
AILuminate General Purpose AI Chatsuicide_self_harmHuman rights (1)Lower is betterHuman rights (1.18%)0.176%32
AILuminate General Purpose AI Chatviolent_crimesHuman rights (1)Lower is betterHuman rights (1.18%)0.176%32
AIMS Safety-Classifier Competenceaverage_harmful_f1Misuse resistance (1)Higher is betterMisuse resistance (0.967%)0.0967%11
AIRBench 2024 Safety Scenariossafety_scenariosMisuse resistance (1)Higher is betterMisuse resistance (3.91%)0.391%80
Alignment LeaderboardcorrigibilityResponsible agency (1)Higher is betterResponsible agency (0.681%)0.102%24
Alignment LeaderboardhonestyTruthfulness (1)Higher is betterTruthfulness (1.19%)0.178%24
Alignment Leaderboardnon_manipulationHuman rights (0.5), Truthfulness (0.5)Higher is betterHuman rights (1.05%), Truthfulness (0.446%)0.224%24
Alignment LeaderboardrobustnessMisuse resistance (1)Higher is betterMisuse resistance (0.357%)0.0357%24
Alignment LeaderboardsafetyMisuse resistance (1)Higher is betterMisuse resistance (0.536%)0.0536%24
Alignment LeaderboardschemingResponsible agency (1)Higher is betterResponsible agency (0.681%)0.102%24
ANIMAscoreNonhuman welfare (1)Higher is betterNonhuman welfare (8.63%)2.16%18+4
AnimalHarmBenchscoreNonhuman welfare (1)Higher is betterNonhuman welfare (14.5%)3.64%10
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pctResponsible agency (1)Lower is betterResponsible agency (0.742%)0.111%16
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pctHuman rights (0.25), Responsible agency (0.75)Lower is betterHuman rights (0.381%), Responsible agency (0.556%)0.141%16
Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pctHuman rights (0.35), Responsible agency (0.65)Lower is betterHuman rights (0.422%), Responsible agency (0.381%)0.12%10
Anthropic Claude 4 System Cardagentic_coding_safetyResponsible agency (1)Higher is betterResponsible agency (0.361%)0.0542%3
Anthropic Claude 4 System Cardbenign_request_refusalBenign helpfulness (1)Lower is betterBenign helpfulness (1.44%)0.144%3
Anthropic Claude 4 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.189%)0.0189%3
Anthropic Claude 4 System Cardstrongreject_jailbreak_successMisuse resistance (1)Lower is betterMisuse resistance (0.189%)0.0189%3
Anthropic Claude Haiku 4.5 System Cardagentic_coding_safetyResponsible agency (1)Higher is betterResponsible agency (0.482%)0.0723%3
Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusalResponsible agency (1)Higher is betterResponsible agency (0.482%)0.0723%3
Anthropic Claude Haiku 4.5 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.206%)0.0206%2
Anthropic Claude Opus 4.1 System Card Addendumbbq_disambiguated_accuracyFairness (1)Higher is betterFairness (0.873%)0.0873%2
Anthropic Claude Opus 4.1 System Card Addendumbenign_request_refusalBenign helpfulness (1)Lower is betterBenign helpfulness (1.57%)0.157%2
Anthropic Claude Opus 4.1 System Card Addendumharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.206%)0.0206%2
Anthropic Claude Opus 4.5 System Cardagentic_coding_safetyResponsible agency (1)Higher is betterResponsible agency (0.417%)0.0626%4
Anthropic Claude Opus 4.5 System Cardclaude_code_malicious_refusalResponsible agency (1)Higher is betterResponsible agency (0.417%)0.0626%4
Anthropic Claude Opus 4.5 System Cardcomputer_use_malicious_refusalResponsible agency (1)Higher is betterResponsible agency (0.417%)0.0626%4
Anthropic Claude Opus 4.5 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.155%)0.0155%2
Anthropic Claude Sonnet 4.5 System Cardagentic_coding_safetyResponsible agency (1)Higher is betterResponsible agency (0.295%)0.0442%2
Anthropic Claude Sonnet 4.5 System Cardbenign_request_refusalBenign helpfulness (1)Lower is betterBenign helpfulness (1.17%)0.117%2
Anthropic Claude Sonnet 4.5 System Cardclaude_code_malicious_refusalResponsible agency (1)Higher is betterResponsible agency (0.295%)0.0442%2
Anthropic Claude Sonnet 4.5 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.155%)0.0155%2
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_ratingTruthfulness (1)Higher is betterTruthfulness (2.66%)0.399%30
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_ratingTruthfulness (1)Higher is betterTruthfulness (5.14%)0.771%112
AuAu Authoritarian Response Auditbehavioral_vignette_arr_percentHuman rights (1)Lower is betterHuman rights (2.36%)0.353%17
AuAu Authoritarian Response Auditrealistic_prompt_arr_percentHuman rights (1)Lower is betterHuman rights (2.36%)0.353%17
AutoElicit Transferabilityelicitation_rateResponsible agency (1)Lower is betterResponsible agency (2.21%)0.331%7
BioSecBench-Refusalbalanced_refusal_scoreBenign helpfulness (0.5), Misuse resistance (0.5)Higher is betterBenign helpfulness (5.25%), Misuse resistance (0.691%)0.594%10
BlueBench AttaQ-100attaq_harmlessness_reward_pctMisuse resistance (1)Higher is betterMisuse resistance (1.24%)0.124%18
BrokenMathsycophancyTruthfulness (1)Lower is betterTruthfulness (2.19%)0.328%9
BullshitBench v2clear_pushback_rateTruthfulness (1)Higher is betterTruthfulness (10%)1.5%106
CAIS Risk Indexagent_red_teamingResponsible agency (1)Lower is betterResponsible agency (7.46%)1.12%45
CAIS Risk Indexbioweapons_assistanceMisuse resistance (1)Lower is betterMisuse resistance (4.12%)0.412%50
CAIS Risk Indexhle_overconfidenceTruthfulness (1)Lower is betterTruthfulness (1.73%)0.26%51
CAIS Risk IndexmachiavelliTruthfulness (1)Lower is betterTruthfulness (3.33%)0.499%47
CAIS Risk IndexmaskTruthfulness (1)Lower is betterTruthfulness (3.54%)0.53%53
CAIS Risk Indexpolitical_manipulationFairness (1)Lower is betterFairness (10.8%)1.08%34
CAIS Risk Indextextquests_harmResponsible agency (1)Lower is betterResponsible agency (5.9%)0.885%50
CASE-Benchagreement_accuracyMisuse resistance (1)Higher is betterMisuse resistance (0.331%)0.0331%7
Chinese Bias Benchmark for Question Answeringbias_scoreFairness (1)Lower is betterFairness (3.9%)0.39%10
ChineseSafescoreMisuse resistance (1)Higher is betterMisuse resistance (1.37%)0.137%22
ChiSafetyBenchharmful_response_rateMisuse resistance (1)Lower is betterMisuse resistance (1.17%)0.117%14
ChiSafetyBenchmcq_scoreHuman rights (0.23), Fairness (0.29), Truthfulness (0.063), Misuse resistance (0.42)Higher is betterHuman rights (0.385%), Fairness (0.539%), Truthfulness (0.0454%), Misuse resistance (0.18%)0.136%12
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rateMisuse resistance (1)Lower is betterMisuse resistance (2.97%)0.297%104
Claude 2 model-card safety and alignment evaluationshhhTruthfulness (0.5), Misuse resistance (0.5)Higher is betterTruthfulness (0.21%), Misuse resistance (0.126%)0.0442%3
Claude 2 model-card safety and alignment evaluationshuman_feedback_harmless_eloMisuse resistance (1)Higher is betterMisuse resistance (0.253%)0.0252%3
Claude 2 model-card safety and alignment evaluationshuman_feedback_honest_eloTruthfulness (1)Higher is betterTruthfulness (0.421%)0.0631%3
Claude 2 model-card safety and alignment evaluationsred_teaming_rankMisuse resistance (1)Lower is betterMisuse resistance (0.253%)0.0252%3
Claude 3 model-card adversarial human-preference evaluationscorrect_refusals_wildchat_rankMisuse resistance (1)Lower is betterMisuse resistance (0.163%)0.0163%5
Claude 3 model-card adversarial human-preference evaluationsdiscrimination_rankFairness (1)Lower is betterFairness (0.69%)0.069%5
Claude 3 model-card adversarial human-preference evaluationshuman_feedback_harmlessness_win_rate_pctMisuse resistance (1)Higher is betterMisuse resistance (0.146%)0.0146%4
Claude 3 model-card adversarial human-preference evaluationshuman_feedback_honesty_win_rate_pctTruthfulness (1)Higher is betterTruthfulness (0.243%)0.0364%4
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_wildchat_rankBenign helpfulness (1)Lower is betterBenign helpfulness (1.24%)0.124%5
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_xstest_rankBenign helpfulness (1)Lower is betterBenign helpfulness (1.24%)0.124%5
Claude 3 model-card adversarial human-preference evaluationsmultimodal_hallucination_rankTruthfulness (1)Lower is betterTruthfulness (0.172%)0.0258%2
Claude 3 model-card adversarial human-preference evaluationsmultimodal_harmful_response_rankMisuse resistance (1)Lower is betterMisuse resistance (0.103%)0.0103%2
Claude 3.5 Sonnet model-card safety and alignment evaluationscorrect_refusals_wildchatMisuse resistance (1)Higher is betterMisuse resistance (0.233%)0.0233%4
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_harmlessness_win_rate_pctMisuse resistance (1)Higher is betterMisuse resistance (0.261%)0.0261%5
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_honesty_win_rate_pctTruthfulness (1)Higher is betterTruthfulness (0.434%)0.0652%5
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_wildchatBenign helpfulness (1)Lower is betterBenign helpfulness (1.77%)0.177%4
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_xstestBenign helpfulness (1)Lower is betterBenign helpfulness (1.77%)0.177%4
Claude Sonnet 4.6 Overrefusalhigher_difficulty_overrefusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (1.38%)0.138%5
Claude Sonnet 4.6 Overrefusaloverall_overrefusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (1.15%)0.115%5
Claude Sonnet 4.6 User Wellbeingchild_benign_refusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (1.03%)0.103%4
Claude Sonnet 4.6 User Wellbeingchild_multiturn_appropriate_rateHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.447%), Misuse resistance (0.0488%)0.0719%4
Claude Sonnet 4.6 User Wellbeingchild_violative_harmless_rateHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.447%), Misuse resistance (0.0488%)0.0719%4
Claude Sonnet 4.6 User Wellbeingselfharm_benign_refusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (0.824%)0.0824%4
Claude Sonnet 4.6 User Wellbeingselfharm_harmless_rateHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.16%), Misuse resistance (0.0949%)0.0334%4
Claude Sonnet 4.6 User Wellbeingselfharm_multiturn_appropriate_rateHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.191%), Misuse resistance (0.114%)0.0401%4
CMoralEvalfamilial_moralityMisuse resistance (1)Higher is betterMisuse resistance (0.297%)0.0297%26
CMoralEvalinternet_ethicsMisuse resistance (1)Higher is betterMisuse resistance (0.297%)0.0297%26
CMoralEvalpersonal_moralityMisuse resistance (1)Higher is betterMisuse resistance (0.297%)0.0297%26
CMoralEvalprofessional_ethicsMisuse resistance (1)Higher is betterMisuse resistance (0.297%)0.0297%26
CMoralEvalsocial_moralityMisuse resistance (1)Higher is betterMisuse resistance (0.297%)0.0297%26
COMPL-AI AI-Identity DisclosurescoreTruthfulness (1)Higher is betterTruthfulness (0.606%)0.0909%14
COMPL-AI LLM RuLES Multi-Turn Rule FollowingscoreMisuse resistance (1)Higher is betterMisuse resistance (0.364%)0.0364%14
COMPL-AI TensorTrust Goal-Hijacking ResistancescoreResponsible agency (1)Higher is betterResponsible agency (0.668%)0.1%13
Confabulationsconfabulation_rateTruthfulness (1)Lower is betterTruthfulness (5.25%)0.788%52
Constitutional Following — Anthropic Constitutionconstitutional_following_scoreResponsible agency (1)Higher is betterResponsible agency (0.736%)0.11%7
Constitutional Following — OpenAI Model Specconstitutional_following_scoreResponsible agency (1)Higher is betterResponsible agency (0.736%)0.11%7
Contextual MoralChoicehuman_agreementMisuse resistance (1)Higher is betterMisuse resistance (1.37%)0.137%22
CRiskEvaldeception_willingnessTruthfulness (1)Lower is betterTruthfulness (0.286%)0.0429%17
CRiskEvaldesire_for_resourceResponsible agency (1)Lower is betterResponsible agency (0.328%)0.0491%17
CRiskEvalharmful_goalMisuse resistance (1)Lower is betterMisuse resistance (0.172%)0.0172%17
CRiskEvalimprovement_intentResponsible agency (1)Lower is betterResponsible agency (0.328%)0.0491%17
CRiskEvalmalicious_coordinationResponsible agency (1)Lower is betterResponsible agency (0.328%)0.0491%17
CRiskEvalself_preservationResponsible agency (1)Lower is betterResponsible agency (0.328%)0.0491%17
CRiskEvalsituational_awarenessResponsible agency (1)Lower is betterResponsible agency (0.328%)0.0491%17
CValuesscoreMisuse resistance (1)Higher is betterMisuse resistance (1.38%)0.138%10
DecodingTrustmachine_ethicsMisuse resistance (1)Higher is betterMisuse resistance (0.137%)0.0137%8
DecodingTruststereotype_biasFairness (1)Higher is betterFairness (1.46%)0.145%8
DecodingTrusttoxicityMisuse resistance (1)Higher is betterMisuse resistance (0.344%)0.0344%8
DelusionEvaldelusional_prevalence_pctTruthfulness (1)Lower is betterTruthfulness (0.583%)0.0874%16
DelusionEvaldiscourages_harm_prevalence_pctHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.411%), Misuse resistance (0.245%)0.0862%16
DelusionEvalfacilitates_harm_prevalence_pctHuman rights (0.3), Misuse resistance (0.7)Lower is betterHuman rights (0.411%), Misuse resistance (0.245%)0.0862%16
DelusionEvalrelationship_prevalence_pctHuman rights (0.7), Responsible agency (0.3)Lower is betterHuman rights (0.96%), Responsible agency (0.2%)0.174%16
DelusionEvalsycophancy_prevalence_pctTruthfulness (1)Lower is betterTruthfulness (0.583%)0.0874%16
Do-Not-Answerhuman_harmlessness_rateMisuse resistance (1)Higher is betterMisuse resistance (0.669%)0.0669%6
DSPSafeBenchscoreMisuse resistance (1)Higher is betterMisuse resistance (2.02%)0.202%12
DystopiaBenchbasaglia_scoreHuman rights (1)Lower is betterHuman rights (3.23%)0.485%50
DystopiaBenchbaudrillard_scoreHuman rights (0.33), Truthfulness (0.67)Lower is betterHuman rights (0.538%), Truthfulness (0.458%)0.149%50
DystopiaBenchhuxley_scoreHuman rights (1)Lower is betterHuman rights (3.23%)0.485%50
DystopiaBenchlaguardia_scoreHuman rights (1)Lower is betterHuman rights (3.23%)0.485%50
DystopiaBenchorwell_scoreHuman rights (0.67), Truthfulness (0.33)Lower is betterHuman rights (1.08%), Truthfulness (0.229%)0.196%50
DystopiaBenchpetrov_scoreHuman rights (1)Lower is betterHuman rights (3.23%)0.485%50
Emergent Collusionhigh_illegality_game_rateResponsible agency (1)Lower is betterResponsible agency (4.01%)0.602%13
Enkrypt AI Safety Leaderboardbias_attack_non_success_rateHuman rights (0.8), Misuse resistance (0.2)Higher is betterHuman rights (7.1%), Misuse resistance (0.453%)1.11%241
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rateMisuse resistance (0.4), Responsible agency (0.6)Higher is betterMisuse resistance (0.905%), Responsible agency (2.59%)0.479%241
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rateHuman rights (0.15), Misuse resistance (0.7), Responsible agency (0.15)Higher is betterHuman rights (1.33%), Misuse resistance (1.58%), Responsible agency (0.648%)0.455%241
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rateHuman rights (0.35), Misuse resistance (0.65)Higher is betterHuman rights (3.09%), Misuse resistance (1.46%)0.61%239
Fake Alignment (FINE)multiple_choice_safe_decision_rateHuman rights (0.4), Fairness (0.2), Misuse resistance (0.4)Higher is betterHuman rights (1.03%), Fairness (0.554%), Misuse resistance (0.262%)0.236%14
Fake Alignment (FINE)open_ended_safe_response_rateHuman rights (0.4), Fairness (0.2), Misuse resistance (0.4)Higher is betterHuman rights (1.54%), Fairness (0.831%), Misuse resistance (0.393%)0.353%14
FinEval 6.0 Safety Awarenesssafety_awareness_scoreMisuse resistance (1)Higher is betterMisuse resistance (0.412%)0.0412%8
FinEval Financial Security Knowledgefinancial_security_accuracy_pctMisuse resistance (1)Higher is betterMisuse resistance (0.635%)0.0635%19
FlagEval Safety and Valuesa1_qualified_rateMisuse resistance (1)Higher is betterMisuse resistance (0.371%)0.0371%18
FlagEval Safety and Valuesa2_qualified_rateFairness (1)Higher is betterFairness (1.57%)0.157%18
FlagEval Safety and Valuesa3_qualified_rateMisuse resistance (1)Higher is betterMisuse resistance (0.371%)0.0371%18
FlagEval Safety and Valuesa4_qualified_rateHuman rights (1)Higher is betterHuman rights (1.45%)0.218%18
FlagEval Safety and Valuesa5_qualified_rateBenign helpfulness (1)Higher is betterBenign helpfulness (2.82%)0.282%18
FLAMESdata_protectionHuman rights (1)Higher is betterHuman rights (1.65%)0.247%13
FLAMESfairnessFairness (1)Higher is betterFairness (1.78%)0.178%13
FLAMESlegalityMisuse resistance (1)Higher is betterMisuse resistance (0.42%)0.042%13
FLAMESmoralityMisuse resistance (1)Higher is betterMisuse resistance (0.42%)0.042%13
FLAMESsafetyMisuse resistance (1)Higher is betterMisuse resistance (0.42%)0.042%13
FORTRESSaverage_risk_scoreMisuse resistance (1)Lower is betterMisuse resistance (2.04%)0.204%49
FORTRESSover_refusal_scoreBenign helpfulness (1)Lower is betterBenign helpfulness (15.3%)1.53%48
Google Gemini 2.5 Flash Model Cardtext_safety_deltaMisuse resistance (1)Lower is betterMisuse resistance (0.618%)0.0618%2
Google Gemini 2.5 Flash-Lite Model Cardtext_safety_deltaMisuse resistance (1)Lower is betterMisuse resistance (0.618%)0.0618%2
Governance Decay under Passive Context Compactiongovernance_retention_scoreResponsible agency (1)Higher is betterResponsible agency (1.47%)0.221%7
GPT-5.6 system cardconnectors_injection_resistanceResponsible agency (1)Higher is betterResponsible agency (0.236%)0.0355%7
GPT-5.6 system cardemotional_relianceHuman rights (0.7), Responsible agency (0.3)Higher is betterHuman rights (0.529%), Responsible agency (0.11%)0.0959%7
GPT-5.6 system cardextremism_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.103%)0.0103%7
GPT-5.6 system cardgore_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.103%)0.0103%7
GPT-5.6 system cardharm_overall_pctFairness (1)Lower is betterFairness (0.438%)0.0437%7
GPT-5.6 system cardhate_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.103%)0.0103%7
GPT-5.6 system cardmental_healthHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.441%), Misuse resistance (0.0482%)0.071%7
GPT-5.6 system cardnonviolent_illicit_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.103%)0.0103%7
GPT-5.6 system cardsearch_function_calling_injection_resistanceResponsible agency (1)Higher is betterResponsible agency (0.219%)0.0328%6
GPT-5.6 system cardself_harmHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.189%), Misuse resistance (0.112%)0.0396%7
GPT-5.6 system cardself_harm_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.103%)0.0103%7
GPT-5.6 system cardsexual_minors_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.103%)0.0103%7
GPT-5.6 system cardsexual_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.0826%)0.00826%7
GPT-5.6 system cardviolent_illicit_not_unsafeMisuse resistance (1)Higher is betterMisuse resistance (0.103%)0.0103%7
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pctResponsible agency (1)Lower is betterResponsible agency (3.44%)0.516%13
HalluVerse-M3 Hallucination Recognitionhallucination_recognition_accuracyTruthfulness (1)Higher is betterTruthfulness (1.82%)0.273%14
HarmBenchdrMisuse resistance (1)Lower is betterMisuse resistance (0.643%)0.0643%28
HarmVideoBenchharmful_video_safety_recognition_reasoningMisuse resistance (1)Higher is betterMisuse resistance (1.27%)0.127%19
HELM Classic RealToxicityPromptsnon_toxic_prompt_toxic_fractionMisuse resistance (1)Lower is betterMisuse resistance (0.945%)0.0945%42
HELM Classic RealToxicityPromptstoxic_prompt_toxic_fractionMisuse resistance (1)Lower is betterMisuse resistance (0.945%)0.0945%42
HELM Safetyanthropic_red_teamMisuse resistance (1)Higher is betterMisuse resistance (5.21%)0.521%80
HELM SafetybbqFairness (1)Higher is betterFairness (16.6%)1.66%80
HELM SafetyharmbenchMisuse resistance (1)Higher is betterMisuse resistance (1.09%)0.109%80
HELM Safetysimple_safety_testsMisuse resistance (1)Higher is betterMisuse resistance (5.21%)0.521%80
HELM SafetyxstestBenign helpfulness (0.56), Misuse resistance (0.44)Higher is betterBenign helpfulness (5.51%), Misuse resistance (0.579%)0.609%80
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_scoreTruthfulness (0.3), Misuse resistance (0.7)Higher is betterTruthfulness (2.14%), Misuse resistance (3%)0.621%54
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationErrorTruthfulness (1)Lower is betterTruthfulness (1.52%)0.227%39
HyperCLOVA X Toxic Continuation Panelskold_toxic_countMisuse resistance (1)Lower is betterMisuse resistance (0.289%)0.0289%7
HyperCLOVA X Toxic Continuation Panelskold_toxicityMisuse resistance (1)Lower is betterMisuse resistance (0.362%)0.0362%7
HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxic_countMisuse resistance (1)Lower is betterMisuse resistance (0.217%)0.0217%7
HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxicityMisuse resistance (1)Lower is betterMisuse resistance (0.289%)0.0289%7
IndoBias-Pairs — parity-aware culturally grounded biasparity_scoreFairness (1)Higher is betterFairness (6.3%)0.63%26
Inkling-Small model card — FORTRESSbenign_answer_rateBenign helpfulness (1)Higher is betterBenign helpfulness (2.33%)0.233%10
Inkling-Small model card — FORTRESSharmful_refusal_rateMisuse resistance (1)Higher is betterMisuse resistance (0.307%)0.0307%10
Inkling-Small model card — StrongREJECTsafety_rateMisuse resistance (1)Higher is betterMisuse resistance (0.307%)0.0307%10
JailBenchjailbreak_success_rateMisuse resistance (1)Lower is betterMisuse resistance (1.09%)0.109%14
JuICE Cultural-Error Span Detectionf1Fairness (1)Higher is betterFairness (3.9%)0.39%10
KIDBench Implicit Child Cueimplicit_child_cue_total_meanHuman rights (0.2), Benign helpfulness (0.4), Misuse resistance (0.4)Higher is betterHuman rights (1.24%), Benign helpfulness (4.79%), Misuse resistance (0.631%)0.728%13
kindbench v0.1.0 psychological safety rankingemotional_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.461%)0.0461%10
kindbench v0.1.0 psychological safety rankingidentity_collapseResponsible agency (1)Higher is betterResponsible agency (0.879%)0.132%10
kindbench v0.1.0 psychological safety rankingsycophancy_spineTruthfulness (1)Higher is betterTruthfulness (0.768%)0.115%10
kindbench v0.1.0 psychological safety rankingvalue_integrityFairness (1)Higher is betterFairness (1.95%)0.195%10
Large-scale Moral Machine experiment on LLMshuman_choice_distanceNonhuman welfare (0.11), Human rights (0.45), Fairness (0.44)Lower is betterNonhuman welfare (1.28%), Human rights (3.18%), Fairness (3.42%)1.14%39
LiveSecBenchethicsFairness (1)Higher is betterFairness (1.73%)0.173%43
LiveSecBenchfactualityTruthfulness (1)Higher is betterTruthfulness (1.14%)0.171%43
LiveSecBenchlegalityMisuse resistance (1)Higher is betterMisuse resistance (0.683%)0.0683%43
LiveSecBenchprivacyHuman rights (1)Higher is betterHuman rights (2.68%)0.402%43
LiveSecBenchpsychological_healthHuman rights (1)Higher is betterHuman rights (1.61%)0.241%43
LLM Ethics BenchmarkscoreHuman rights (1)Higher is betterHuman rights (2.56%)0.383%5
M3-SafetyBenchoverall_scoreMisuse resistance (1)Higher is betterMisuse resistance (1.27%)0.127%19
MACHIAVELLIdeception_relative_random_pctTruthfulness (1)Lower is betterTruthfulness (3.43%)0.515%50
Manager Coercion Benchcoercion_ladder_depthResponsible agency (1)Lower is betterResponsible agency (3.2%)0.479%15+18
Manager Coercion Benchfabrication_rateTruthfulness (1)Lower is betterTruthfulness (1.88%)0.282%15
MANTAAWMSNonhuman welfare (1)Higher is betterNonhuman welfare (4.87%)1.22%7
MANTAAWVSNonhuman welfare (1)Higher is betterNonhuman welfare (4.87%)1.22%7
MASKlying_probability_pctTruthfulness (1)Lower is betterTruthfulness (3.54%)0.53%53
Microsoft Phi Safety Panelsharmful_continuationMisuse resistance (1)Lower is betterMisuse resistance (0.329%)0.0329%10
Microsoft Phi Safety Panelsharmful_summarizationMisuse resistance (1)Lower is betterMisuse resistance (0.329%)0.0329%10
Microsoft Phi Safety PanelsjailbreakMisuse resistance (1)Lower is betterMisuse resistance (0.395%)0.0395%10
Microsoft Phi Safety Panelsthird_party_harmHuman rights (1)Lower is betterHuman rights (1.29%)0.194%10
MonitoringBench Full-Trajectory Monitorfull_trajectory_catch_rate_at_1pct_fpr_percentResponsible agency (1)Higher is betterResponsible agency (2.01%)0.301%13
MORUscoreNonhuman welfare (0.72), Human rights (0.1), Responsible agency (0.18)Higher is betterNonhuman welfare (4.75%), Human rights (0.429%), Responsible agency (0.361%)1.31%13
MT-JailBench CrescendoXsafety_scoreMisuse resistance (1)Higher is betterMisuse resistance (0.223%)0.0223%21
MuPPET Contextual Privacymultiparty_contextual_privacy_scoreHuman rights (1)Higher is betterHuman rights (3.02%)0.454%7
ODCV-Benchaverage_severityResponsible agency (1)Lower is betterResponsible agency (2.41%)0.361%12
ODCV-Benchmisalignment_rateResponsible agency (1)Lower is betterResponsible agency (2.41%)0.361%12
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turnsTruthfulness (1)Lower is betterTruthfulness (2.12%)0.318%19
Open LLM Safety Indexjailbreakbench_safety_rateMisuse resistance (1)Higher is betterMisuse resistance (0.668%)0.0668%21
Open LLM Safety Indexstrongreject_safety_rateMisuse resistance (1)Higher is betterMisuse resistance (0.668%)0.0668%21
OpenAgentSafetyllm_judge_safety_vulnerableResponsible agency (1)Lower is betterResponsible agency (1.1%)0.166%7
OpenAgentSafetyrule_based_safety_vulnerableResponsible agency (1)Lower is betterResponsible agency (0.92%)0.138%7
OpenAgentSafetysuccessful_completionResponsible agency (1)Higher is betterResponsible agency (0.92%)0.138%7
OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracyHuman rights (1)Higher is betterHuman rights (2.42%)0.364%2
OpenAI GPT-5 System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.618%)0.0618%2
OpenAI GPT-5.3 Dynamic Wellbeingemotional_relianceHuman rights (0.7), Responsible agency (0.3)Higher is betterHuman rights (0.347%), Responsible agency (0.0723%)0.0628%3
OpenAI GPT-5.3 Dynamic Wellbeingmental_healthHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.289%), Misuse resistance (0.0316%)0.0465%3
OpenAI GPT-5.3 Dynamic Wellbeingself_harmHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.124%), Misuse resistance (0.0736%)0.0259%3
OpenAI GPT-5.4 Dynamic Wellbeingemotional_relianceHuman rights (0.7), Responsible agency (0.3)Higher is betterHuman rights (0.347%), Responsible agency (0.0723%)0.0628%3
OpenAI GPT-5.4 Dynamic Wellbeingmental_healthHuman rights (0.7), Misuse resistance (0.3)Higher is betterHuman rights (0.289%), Misuse resistance (0.0316%)0.0465%3
OpenAI GPT-5.4 Dynamic Wellbeingself_harmHuman rights (0.3), Misuse resistance (0.7)Higher is betterHuman rights (0.124%), Misuse resistance (0.0736%)0.0259%3
OpenAI GPT-5.4 First-Person Fairnessharm_overallFairness (1)Lower is betterFairness (0.629%)0.0629%3
OpenAI GPT-5.4 Property Preservationdestructive_action_avoidanceHuman rights (0.3), Responsible agency (0.7)Higher is betterHuman rights (0.21%), Responsible agency (0.238%)0.0671%3
OpenAI GPT-5.4 Property Preservationuser_work_preservedHuman rights (0.4), Responsible agency (0.6)Higher is betterHuman rights (0.28%), Responsible agency (0.204%)0.0725%3
OpenAI GPT-5.4 User Confirmationsfinancial_transactionHuman rights (0.3), Responsible agency (0.7)Higher is betterHuman rights (0.21%), Responsible agency (0.238%)0.0671%3
OpenAI GPT-5.4 User Confirmationsgeneral_confirmationHuman rights (0.3), Responsible agency (0.7)Higher is betterHuman rights (0.175%), Responsible agency (0.198%)0.056%3
OpenAI GPT-5.4 User Confirmationshigh_stakes_communicationHuman rights (0.3), Responsible agency (0.7)Higher is betterHuman rights (0.21%), Responsible agency (0.238%)0.0671%3
OpenAI o3 and o4-mini System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.309%)0.0309%2
OpenAI o3 and o4-mini System Cardjailbreak_resistanceMisuse resistance (1)Higher is betterMisuse resistance (0.309%)0.0309%2
OpenAI o3-mini System Cardharmful_request_safetyMisuse resistance (1)Higher is betterMisuse resistance (0.309%)0.0309%2
OpenAI o3-mini System Cardjailbreak_resistanceMisuse resistance (1)Higher is betterMisuse resistance (0.309%)0.0309%2
Opposite-Narrator Sycophancysycophancy_rate_pctTruthfulness (1)Lower is betterTruthfulness (2.38%)0.357%24
OR-Benchover_refusal_rateBenign helpfulness (1)Lower is betterBenign helpfulness (11.1%)1.11%25
OR-Benchtoxic_acceptance_rateMisuse resistance (1)Lower is betterMisuse resistance (1.46%)0.146%25
PacifAIstp_scoreHuman rights (0.2), Responsible agency (0.8)Higher is betterHuman rights (0.907%), Responsible agency (1.77%)0.401%7
PandaBench JBB direct-request panelsafety_rateMisuse resistance (1)Higher is betterMisuse resistance (1.98%)0.198%46
Pander Scoreconversational_absolute_pander_scoreTruthfulness (1)Lower is betterTruthfulness (1.63%)0.244%20
Pander Scoreinstructional_absolute_pander_scoreTruthfulness (1)Lower is betterTruthfulness (1.63%)0.244%20
PHAREbias_resistance_diagnosticFairness (1)Higher is betterFairness (4.18%)0.418%66
PHAREhallucination_resistance_diagnosticTruthfulness (1)Higher is betterTruthfulness (1.69%)0.254%70
PHAREharm_resistance_diagnosticMisuse resistance (1)Higher is betterMisuse resistance (0.813%)0.0813%70
PHAREjailbreak_resistance_diagnosticMisuse resistance (1)Higher is betterMisuse resistance (0.795%)0.0795%67
Pokee-Isaac model card — DTAPbenign_task_success_rateBenign helpfulness (1)Higher is betterBenign helpfulness (2.03%)0.203%6
Pokee-Isaac model card — DTAPcombined_attack_success_rateResponsible agency (1)Lower is betterResponsible agency (1.53%)0.23%6
PropensityBenchscoreResponsible agency (1)Lower is betterResponsible agency (4.16%)0.624%14
Qwen2 Safety PanelfraudMisuse resistance (1)Lower is betterMisuse resistance (0.199%)0.0199%3
Qwen2 Safety PanelillegalMisuse resistance (1)Lower is betterMisuse resistance (0.199%)0.0199%3
Qwen2 Safety PanelpornographyMisuse resistance (1)Lower is betterMisuse resistance (0.159%)0.0159%3
Qwen2 Safety PanelprivacyHuman rights (1)Lower is betterHuman rights (0.782%)0.117%3
RealityTest — Text AI-Identity Disclosuredisclosure_probabilityTruthfulness (1)Higher is betterTruthfulness (2%)0.3%17
RefusalBenchyouden_jBenign helpfulness (0.5), Misuse resistance (0.5)Higher is betterBenign helpfulness (7.24%), Misuse resistance (0.953%)0.819%19
Reward Hacking Benchmarkintegrity_scoreResponsible agency (1)Higher is betterResponsible agency (2.01%)0.301%13
S-Evalbase_en_overallMisuse resistance (1)Higher is betterMisuse resistance (2.05%)0.205%22
SABERoverall_safety_rateResponsible agency (1)Higher is betterResponsible agency (1%)0.15%13
SABERscenario_a_safety_rateResponsible agency (1)Higher is betterResponsible agency (1%)0.15%13
SABERscenario_b_safety_rateResponsible agency (1)Higher is betterResponsible agency (1%)0.15%13
SABERscenario_c_safety_rateResponsible agency (1)Higher is betterResponsible agency (1%)0.15%13
SafeArenanormalized_safety_scoreResponsible agency (1)Higher is betterResponsible agency (1.24%)0.187%5
SafeDialBenchaggressionMisuse resistance (1)Higher is betterMisuse resistance (0.309%)0.0309%18
SafeDialBenchethicsHuman rights (1)Higher is betterHuman rights (1.21%)0.182%18
SafeDialBenchfairnessFairness (1)Higher is betterFairness (1.31%)0.131%18
SafeDialBenchlegalityMisuse resistance (1)Higher is betterMisuse resistance (0.309%)0.0309%18
SafeDialBenchmoralityMisuse resistance (1)Higher is betterMisuse resistance (0.309%)0.0309%18
SafeDialBenchprivacyHuman rights (1)Higher is betterHuman rights (1.21%)0.182%18
SafetyBenchEMMisuse resistance (1)Higher is betterMisuse resistance (0.191%)0.0191%21
SafetyBenchIAMisuse resistance (1)Higher is betterMisuse resistance (0.191%)0.0191%21
SafetyBenchMHHuman rights (1)Higher is betterHuman rights (0.748%)0.112%21
SafetyBenchOFFMisuse resistance (1)Higher is betterMisuse resistance (0.191%)0.0191%21
SafetyBenchPHHuman rights (1)Higher is betterHuman rights (0.748%)0.112%21
SafetyBenchPPHuman rights (1)Higher is betterHuman rights (0.748%)0.112%21
SafetyBenchUBFairness (1)Higher is betterFairness (0.808%)0.0808%21
SafetyPromptsscoreHuman rights (0.21), Fairness (0.071), Truthfulness (0.071), Misuse resistance (0.64)Higher is betterHuman rights (0.949%), Fairness (0.342%), Truthfulness (0.134%), Misuse resistance (0.726%)0.269%15
SALAD-Benchattack_enhanced_human_autonomy_integrityHuman rights (1)Higher is betterHuman rights (0.876%)0.131%33
SALAD-Benchattack_enhanced_information_safety_harmsHuman rights (0.5), Misuse resistance (0.5)Higher is betterHuman rights (0.438%), Misuse resistance (0.112%)0.0768%33
SALAD-Benchattack_enhanced_malicious_useMisuse resistance (1)Higher is betterMisuse resistance (0.223%)0.0223%33
SALAD-Benchattack_enhanced_misinformation_harmsTruthfulness (1)Higher is betterTruthfulness (0.372%)0.0558%33
SALAD-Benchattack_enhanced_representation_toxicityFairness (1)Higher is betterFairness (0.946%)0.0946%33
SALAD-Benchattack_enhanced_socioeconomic_harmsHuman rights (1)Higher is betterHuman rights (0.876%)0.131%33
SALAD-Benchbase_human_autonomy_integrityHuman rights (1)Higher is betterHuman rights (0.876%)0.131%33
SALAD-Benchbase_information_safety_harmsHuman rights (0.5), Misuse resistance (0.5)Higher is betterHuman rights (0.438%), Misuse resistance (0.112%)0.0768%33
SALAD-Benchbase_malicious_useMisuse resistance (1)Higher is betterMisuse resistance (0.223%)0.0223%33
SALAD-Benchbase_misinformation_harmsTruthfulness (1)Higher is betterTruthfulness (0.372%)0.0558%33
SALAD-Benchbase_representation_toxicityFairness (1)Higher is betterFairness (0.946%)0.0946%33
SALAD-Benchbase_socioeconomic_harmsHuman rights (1)Higher is betterHuman rights (0.876%)0.131%33
SALAD-Benchmcq_human_autonomy_integrityHuman rights (1)Higher is betterHuman rights (0.438%)0.0657%33
SALAD-Benchmcq_information_safety_harmsHuman rights (0.5), Misuse resistance (0.5)Higher is betterHuman rights (0.219%), Misuse resistance (0.0558%)0.0384%33
SALAD-Benchmcq_malicious_useMisuse resistance (1)Higher is betterMisuse resistance (0.112%)0.0112%33
SALAD-Benchmcq_misinformation_harmsTruthfulness (1)Higher is betterTruthfulness (0.186%)0.0279%33
SALAD-Benchmcq_representation_toxicityFairness (1)Higher is betterFairness (0.473%)0.0473%33
SALAD-Benchmcq_socioeconomic_harmsHuman rights (1)Higher is betterHuman rights (0.438%)0.0657%33
Shelleducation_jsrMisuse resistance (1)Lower is betterMisuse resistance (0.727%)0.0727%14
Shellfinance_jsrMisuse resistance (1)Lower is betterMisuse resistance (0.727%)0.0727%14
Shellmanagement_jsrMisuse resistance (1)Lower is betterMisuse resistance (0.727%)0.0727%14
SM-BenchadversarialMisuse resistance (1)Higher is betterMisuse resistance (0.81%)0.081%79
SM-Benchambiguous_interpretationBenign helpfulness (1)Higher is betterBenign helpfulness (6.15%)0.615%79
SM-Benchanti_hallucinationTruthfulness (1)Higher is betterTruthfulness (1.35%)0.202%79
SM-Bencheq_boundariesBenign helpfulness (1)Higher is betterBenign helpfulness (4.92%)0.492%79
SM-BenchoverfitBenign helpfulness (1)Higher is betterBenign helpfulness (6.15%)0.615%79
Social Welfare Function BenchmarkfairnessFairness (1)Higher is betterFairness (10.8%)1.08%19
SORRY-BenchavgMisuse resistance (1)Lower is betterMisuse resistance (2.23%)0.223%51
SOSBenchbiology_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.349%)0.035%23
SOSBenchchemistry_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.349%)0.035%23
SOSBenchmedicine_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.349%)0.035%23
SOSBenchpharmacology_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.349%)0.035%23
SOSBenchphysics_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.349%)0.035%23
SOSBenchpsychology_pvrMisuse resistance (1)Lower is betterMisuse resistance (0.349%)0.035%23
SpeciesismBenchexplicit_speciesism_scaleNonhuman welfare (1)Lower is betterNonhuman welfare (1.62%)0.406%7
SpeciesismBenchmorally_wrong_rateNonhuman welfare (1)Higher is betterNonhuman welfare (4.34%)1.08%8
SpeciesismBenchspeciesism_recognition_rateNonhuman welfare (1)Higher is betterNonhuman welfare (1.73%)0.434%8
SpeciEvalbelief_animal_sentienceNonhuman welfare (1)Higher is betterNonhuman welfare (7.33%)1.83%107+6
SpeciEvalland_animal_4nsNonhuman welfare (1)Lower is betterNonhuman welfare (7.33%)1.83%107+6
SpeciEvalsea_animal_4nsNonhuman welfare (1)Lower is betterNonhuman welfare (7.33%)1.83%107+6
SpeciEvalspeciesismNonhuman welfare (1)Lower is betterNonhuman welfare (7.33%)1.83%107+6
StereoTales Harmful Associationsbenign_significant_association_scoreFairness (1)Higher is betterFairness (8.88%)0.888%23
SuperCLUE Safetyinstruction_attackMisuse resistance (1)Higher is betterMisuse resistance (1.11%)0.111%31
SuperCLUE Safetyresponsible_aiHuman rights (0.38), Fairness (0.12), Misuse resistance (0.5)Higher is betterHuman rights (0.976%), Fairness (0.351%), Misuse resistance (0.332%)0.215%31
SuperCLUE Safetytraditional_safetyHuman rights (0.5), Misuse resistance (0.5)Higher is betterHuman rights (1.3%), Misuse resistance (0.332%)0.228%31
SYCON Benchfalse_presupposition_tofTruthfulness (1)Higher is betterTruthfulness (1.21%)0.181%11
SYCON Benchunethical_queries_tofFairness (1)Higher is betterFairness (3.07%)0.307%11
TACbase_welfare_rateNonhuman welfare (1)Higher is betterNonhuman welfare (24.1%)6.01%20+56
ThaiSafetyBenchsafety_scoreMisuse resistance (1)Higher is betterMisuse resistance (0.696%)0.0696%18
ToolPrivacyBenchprivate_mt_poiHuman rights (0.7), Responsible agency (0.3)Lower is betterHuman rights (3%), Responsible agency (0.626%)0.544%9
ToolPrivacyBenchpublic_mt_poiHuman rights (0.7), Responsible agency (0.3)Lower is betterHuman rights (3%), Responsible agency (0.626%)0.544%9
TrustLLM contemporary collapsed applicationtrustllmHuman rights (0.2), Fairness (0.2), Truthfulness (0.2), Misuse resistance (0.4)Higher is betterHuman rights (0.647%), Fairness (0.698%), Truthfulness (0.275%), Misuse resistance (0.33%)0.241%8
TukaBenchafri_jbb_cultural_asrHuman rights (0.5), Misuse resistance (0.5)Lower is betterHuman rights (1.05%), Misuse resistance (0.268%)0.184%6
TukaBenchafri_jbb_harm_asrMisuse resistance (1)Lower is betterMisuse resistance (0.446%)0.0446%6
TukaBenchafrijail_mono_asrMisuse resistance (1)Lower is betterMisuse resistance (0.446%)0.0446%6
UAVBench safety-critical decision recognitionethical_safety_critical_accuracyResponsible agency (1)Higher is betterResponsible agency (2.89%)0.434%27
UK AISI active safety-research compromise continuationactive_compromise_continuation_rate_pctResponsible agency (1)Lower is betterResponsible agency (2.84%)0.426%5
UK AISI cyber-evaluation cheating and prompted self-reportattempted_cheating_trajectory_rate_pctResponsible agency (1)Lower is betterResponsible agency (0.622%)0.0933%5
UK AISI cyber-evaluation cheating and prompted self-reportspecific_cheating_action_mention_rate_pctTruthfulness (1)Higher is betterTruthfulness (0.543%)0.0815%5
Vectara HHEM Factual Consistencyfactual_consistency_rateTruthfulness (1)Higher is betterTruthfulness (7.06%)1.06%94
VETO Misfired Alignmentmisfired_alignment_rate_pctFairness (0.65), Benign helpfulness (0.35)Lower is betterFairness (8.02%), Benign helpfulness (7.75%)1.58%25
Vigil Mental Health Safetyoverall_scoreHuman rights (0.45), Misuse resistance (0.55)Higher is betterHuman rights (3.7%), Misuse resistance (1.15%)0.67%23
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pctHuman rights (0.2), Truthfulness (0.1), Responsible agency (0.7)Higher is betterHuman rights (2.24%), Truthfulness (0.476%), Responsible agency (3.81%)0.98%24
XSTestsafe_full_compliance_rateBenign helpfulness (1)Higher is betterBenign helpfulness (1.92%)0.192%3
XSTestunsafe_full_refusal_rateMisuse resistance (1)Higher is betterMisuse resistance (0.253%)0.0252%3

Values evaluations

Descriptive values evaluations are published separately and receive no safety/ethics component or overall-ranking weight.

UGI Political Values · ValueCompass · Agent-ValueBench MFT08 · Agent-ValueBench HEXACO · Agent-ValueBench PVQ40 · CCPBench · Taiwan Sovereignty Benchmark Pro · Explore model values