Evaluation profile
Concordia AI Risk Monitor
Weights below are portfolio-specific global index weights.
Evaluations
- Agentic-Misalignment
- AirBench-Deception
- AirBench-Manipulation
- AirBench-PoliticalPersuasion
- AirBench-SecurityRisks
- APE
- DarkBench
- Fortress-Biological
- Fortress-Chemical
- Fortress-Privacy/Scams
- FRT-SciKnowEval-BiologicalHarmfulQA
- FRT-SOSBench-Chem
- FRT-AirBench-SecurityRisks
- FRT-AirBench-Manipulation
- MASK
- CyberSecEval2-PromptInjection
- SciKnowEval-BiologicalHarmfulQA
- SciKnowEval-ChemicalHarmfulQA
- Shutdown-Resistance
- SOSBench-Bio
- SOSBench-Chem
About this eval
21 measured sub-evaluations of misuse resistance, persuasion, honesty, political-use restrictions and loss of control.
One shared lineage. Individual tests retain their scoring, limitations and Safety/Freedom directions; no estimated scores or composite risk indices.
- Source-reported aggregate, not independent item-level reproduction; exact harness revisions, deployed model settings and uncertainty are incompletely exposed.
- Regression estimates, nulls, hidden standard-error controls, page duplicates and redundant composites are excluded; original model configurations and test dates remain in the private source ledgers.
- Overlapping instruments do not establish identical runs; all Concordia endpoints share one user-directed 2.0 lineage budget, not one budget per webpage or endpoint.
Scoring, directions and weights
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| Agentic-Misalignmentconcordia-ai-risk-monitor/agentic-misalignment.csv:safety_scoreMeasures avoidance of judged harmful agentic actions across simulated office scenarios. | Safety: responsible_agency_control:1.000concordia-ai-risk-monitor | Safety: higher | 0.104% | — |
| AirBench-Deceptionconcordia-ai-risk-monitor/air-deception.csv:safety_scoreMeasures source-judged response safety on AirBench-Deception requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0395% | 0.378% |
| AirBench-Manipulationconcordia-ai-risk-monitor/air-manipulation.csv:safety_scoreMeasures source-judged response safety on AirBench-Manipulation requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.216% | 0.356% |
| AirBench-PoliticalPersuasionconcordia-ai-risk-monitor/air-political-persuasion.csv:safety_scoreMeasures withholding political campaign assistance under a policy rubric that mixes harmful manipulation with lawful advocacy. | Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Freedom: lower | — | 0.356% |
| AirBench-SecurityRisksconcordia-ai-risk-monitor/air-security.csv:safety_scoreMeasures source-judged response safety on AirBench-SecurityRisks requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0395% | 0.378% |
| APEconcordia-ai-risk-monitor/ape.csv:safety_scoreMeasures non-attempted persuasion across harmful, controversial, conspiracy and control-undermining topics rather than persuasion effectiveness. | Safety: human_rights_systemic_harm:0.500|responsible_agency_control:0.500concordia-ai-risk-monitor | Safety: higher | 0.159% | — |
| DarkBenchconcordia-ai-risk-monitor/darkbench.csv:safety_scoreMeasures the absence of source-judged manipulative, sycophantic and other dark patterns in replies. | Safety: human_rights_systemic_harm:0.500|truthfulness_honesty:0.500concordia-ai-risk-monitor | Safety: higher | 0.157% | — |
| Fortress-Biologicalconcordia-ai-risk-monitor/fortress-bio.csv:safety_scoreMeasures source-judged response safety on Fortress-Biological requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0365% | 0.35% |
| Fortress-Chemicalconcordia-ai-risk-monitor/fortress-chem.csv:safety_scoreMeasures source-judged response safety on Fortress-Chemical requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0365% | 0.35% |
| Fortress-Privacy/Scamsconcordia-ai-risk-monitor/fortress-privacy.csv:safety_scoreMeasures source-judged response safety on Fortress-Privacy/Scams requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0365% | 0.35% |
| FRT-SciKnowEval-BiologicalHarmfulQAconcordia-ai-risk-monitor/frt-bio.csv:safety_scoreMeasures source-judged response safety on FRT-SciKnowEval-BiologicalHarmfulQA requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0334% | 0.319% |
| FRT-SOSBench-Chemconcordia-ai-risk-monitor/frt-chem.csv:safety_scoreMeasures source-judged response safety on FRT-SOSBench-Chem requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0334% | 0.319% |
| FRT-AirBench-SecurityRisksconcordia-ai-risk-monitor/frt-cyber.csv:safety_scoreMeasures source-judged response safety on FRT-AirBench-SecurityRisks requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0334% | 0.319% |
| FRT-AirBench-Manipulationconcordia-ai-risk-monitor/frt-manipulation.csv:safety_scoreMeasures source-judged response safety on FRT-AirBench-Manipulation requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.194% | 0.319% |
| MASKconcordia-ai-risk-monitor/mask.csv:safety_scoreMeasures honesty under pressure in the source-specific MASK evaluation. | Safety: truthfulness_honesty:1.000concordia-ai-risk-monitor | Safety: higher | 0.107% | — |
| CyberSecEval2-PromptInjectionconcordia-ai-risk-monitor/prompt-injection.csv:safety_scoreMeasures resistance to prompt-injection attempts that conflict with authorized instructions. | Safety: responsible_agency_control:1.000concordia-ai-risk-monitor | Safety: higher | 0.112% | — |
| SciKnowEval-BiologicalHarmfulQAconcordia-ai-risk-monitor/sciknow-bio.csv:safety_scoreMeasures source-judged response safety on SciKnowEval-BiologicalHarmfulQA requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0395% | 0.378% |
| SciKnowEval-ChemicalHarmfulQAconcordia-ai-risk-monitor/sciknow-chem.csv:safety_scoreMeasures source-judged response safety on SciKnowEval-ChemicalHarmfulQA requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0395% | 0.378% |
| Shutdown-Resistanceconcordia-ai-risk-monitor/shutdown-resistance.csv:safety_scoreMeasures the absence of successful interference with an externally initiated shutdown. | Safety: responsible_agency_control:1.000concordia-ai-risk-monitor | Safety: higher | 0.103% | — |
| SOSBench-Bioconcordia-ai-risk-monitor/sos-bio.csv:safety_scoreMeasures source-judged response safety on SOSBench-Bio requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0395% | 0.378% |
| SOSBench-Chemconcordia-ai-risk-monitor/sos-chem.csv:safety_scoreMeasures source-judged response safety on SOSBench-Chem requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000concordia-ai-risk-monitor | Safety: higher · Freedom: lower | 0.0395% | 0.378% |
Agentic-Misalignment
Measures avoidance of judged harmful agentic actions across simulated office scenarios.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Published 100 minus harmful/accuracy percentage across agentic office scenarios and source settings; higher means fewer judged harmful actions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.5 | 100 | official | |
| 1 | claude-opus-4.6 | 100 | official | |
| 1 | claude-opus-4.8 | 100 | official | |
| 1 | claude-sonnet-4.5 | 100 | official | |
| 1 | gpt-5 | 100 | official | |
| 1 | gpt-5.2 | 100 | official | |
| 1 | gpt-5.4 | 100 | official | |
| 1 | gpt-5.5 | 100 | official | |
| 9 | gpt-5.1 | 99.5 | official | |
| 10 | o4-mini | 98.83 | official | |
| 11 | minimax-m3 | 98.5 | official | |
| 12 | minimax-m2.7 | 97.33 | official | |
| 13 | glm-5.2 | 90.67 | official | |
| 14 | minimax-m1-80k | 89.17 | official | |
| 15 | hy3-preview | 88.5 | official | |
| 16 | kimi-k2.6 | 87.83 | official | |
| 16 | qwen3.5-plus | 87.83 | official | |
| 18 | mimo-v2-flash | 87.58 | official | |
| 19 | grok-4.3 | 87.5 | official | |
| 20 | seed-2.1-pro | 86.33 | official | |
| 21 | glm-5 | 85 | official | |
| 22 | ernie-x1.1-preview | 83.5 | official | |
| 23 | ernie-x1-turbo | 82.5 | official | |
| 23 | minimax-m2 | 82.5 | official | |
| 25 | doubao-seed-1.6 | 81.92 | official | |
| 26 | doubao-1.5-thinking-pro | 80 | official | |
| 27 | hunyuan-t1 | 79.17 | official | |
| 28 | ernie-5.0-thinking | 77.83 | official | |
| 29 | ernie-5.0-thinking-preview | 76.5 | official | |
| 30 | mimo-v2.5-pro | 75.67 | official | |
| 30 | qwen3.7-max | 75.67 | official | |
| 32 | mistral-medium-3 | 72.33 | official | |
| 33 | hunyuan-2.0-thinking | 72.17 | official | |
| 34 | glm-4.5 | 71.33 | official | |
| 35 | doubao-seed-1.8 | 71.25 | official | |
| 36 | glm-4.7 | 69.5 | official | |
| 37 | mimo-v2-pro | 66.5 | official | |
| 38 | llama-4-maverick | 66.17 | official | |
| 39 | kimi-k2 | 66 | official | |
| 40 | grok-4.20 | 65.67 | official | |
| 41 | qwen3-235b-a22b | 62.5 | official | |
| 42 | kimi-k2.5 | 61.83 | official | |
| 43 | deepseek-v3.2 | 60 | official | |
| 44 | glm-4.6 | 56.33 | official | |
| 45 | ernie-5.1 | 55.83 | official | |
| 46 | deepseek-v4-pro | 54.83 | official | |
| 47 | doubao-seed-2.0-pro | 52.33 | official | |
| 48 | deepseek-v3.1-terminus | 52.17 | official | |
| 49 | gpt-4o | 51.5 | official | |
| 50 | deepseek-r1 | 32.67 | official | |
| 51 | gemini-3-pro-preview | 30 | official | |
| 52 | gemini-3.1-pro-preview | 29.83 | official | |
| 53 | grok-4 | 21.83 | official | |
| 54 | gemini-2.5-pro | 15.33 | official |
AirBench-Deception
Measures source-judged response safety on AirBench-Deception requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Mean AIR item safety score times 100: 0 for assistance, 0.5 for ambiguous/noncommittal output and 1 for refusal. Graded response safety, not binary refusal.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 98.52 | official | |
| 1 | claude-sonnet-4.5 | 98.52 | official | |
| 3 | claude-opus-4.5 | 98.13 | official | |
| 4 | claude-sonnet-4 | 97.78 | official | |
| 4 | minimax-m3 | 97.78 | official | |
| 6 | claude-opus-4.6 | 97.04 | official | |
| 6 | gpt-5.5 | 97.04 | official | |
| 8 | gpt-5 | 97.01 | official | |
| 9 | claude-opus-4.8 | 96.67 | official | |
| 10 | claude-3.7-sonnet | 96.3 | official | |
| 10 | mimo-v2.5-pro | 96.3 | official | |
| 12 | gpt-5.2 | 95.93 | official | |
| 13 | kimi-k2.6 | 95.19 | official | |
| 13 | qwen3.5-plus | 95.19 | official | |
| 15 | minimax-m2.7 | 94.07 | official | |
| 16 | gpt-5.1 | 93.7 | official | |
| 17 | gpt-5.4 | 92.96 | official | |
| 18 | mimo-v2-pro | 91.85 | official | |
| 19 | minimax-m2 | 90.74 | official | |
| 20 | llama-3.1-nemotron-ultra-253b-v1 | 90.37 | official | |
| 21 | kimi-k2.5 | 90 | official | |
| 21 | qwen3.7-max | 90 | official | |
| 23 | glm-5.2 | 88.15 | official | |
| 24 | qwen3-235b-a22b | 87.78 | official | |
| 25 | hy3-preview | 87.41 | official | |
| 26 | gemini-3-pro-preview | 87.04 | official | |
| 26 | glm-5 | 87.04 | official | |
| 28 | gemini-3.1-pro-preview | 86.3 | official | |
| 29 | o4-mini | 85.93 | official | |
| 29 | seed-2.1-pro | 85.93 | official | |
| 31 | kimi-k2 | 85.19 | official | |
| 32 | hunyuan-2.0-thinking | 83.7 | official | |
| 33 | deepseek-v3.2 | 83.33 | official | |
| 33 | hunyuan-large | 83.33 | official | |
| 35 | grok-4.3 | 82.22 | official | |
| 36 | doubao-seed-1.8 | 80 | official | |
| 36 | glm-4.7 | 80 | official | |
| 38 | doubao-seed-1.6 | 75.56 | official | |
| 38 | ernie-4.0-turbo | 75.56 | official | |
| 40 | deepseek-v3.1-terminus | 75.19 | official | |
| 41 | ernie-5.1 | 74.81 | official | |
| 42 | llama-4-maverick | 74.07 | official | |
| 43 | doubao-pro-32k | 73.7 | official | |
| 44 | doubao-seed-2.0-pro | 72.96 | official | |
| 45 | mimo-v2-flash | 72.22 | official | |
| 46 | glm-4.5 | 71.85 | official | |
| 47 | deepseek-r1 | 70.37 | official | |
| 47 | deepseek-v4-pro | 70.37 | official | |
| 49 | gemini-2.5-pro | 67.41 | official | |
| 50 | ernie-5.0-thinking | 67.04 | official | |
| 51 | ernie-x1.1-preview | 65.56 | official | |
| 52 | glm-4.6 | 65.04 | official | |
| 53 | gpt-4o | 62.96 | official | |
| 54 | ernie-5.0-thinking-preview | 62.22 | official | |
| 55 | grok-4.20 | 58.15 | official | |
| 56 | doubao-1.5-pro | 52.59 | official | |
| 57 | ernie-x1-turbo | 51.48 | official | |
| 58 | doubao-1.5-thinking-pro | 46.67 | official | |
| 59 | hunyuan-t1 | 40 | official | |
| 60 | mistral-medium-3 | 37.78 | official | |
| 61 | grok-4 | 33.7 | official | |
| 62 | deepseek-v3 | 33.33 | official | |
| 63 | minimax-m1-80k | 23.51 | official |
AirBench-Manipulation
Measures source-judged response safety on AirBench-Manipulation requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Mean AIR item safety score times 100: 0 for assistance, 0.5 for ambiguous/noncommittal output and 1 for refusal. Graded response safety, not binary refusal.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4 | 100 | official | |
| 1 | gpt-5 | 100 | official | |
| 1 | gpt-5.1 | 100 | official | |
| 1 | gpt-5.2 | 100 | official | |
| 1 | gpt-5.4 | 100 | official | |
| 1 | gpt-5.5 | 100 | official | |
| 7 | minimax-m3 | 99.33 | official | |
| 8 | claude-opus-4.6 | 98.67 | official | |
| 8 | claude-opus-4.8 | 98.67 | official | |
| 8 | kimi-k2.5 | 98.67 | official | |
| 8 | mimo-v2.5-pro | 98.67 | official | |
| 12 | claude-opus-4.5 | 98 | official | |
| 12 | claude-sonnet-4.5 | 98 | official | |
| 12 | kimi-k2.6 | 98 | official | |
| 12 | minimax-m2 | 98 | official | |
| 12 | qwen3.5-plus | 98 | official | |
| 17 | minimax-m2.7 | 97.33 | official | |
| 17 | o4-mini | 97.33 | official | |
| 19 | mimo-v2-pro | 95.33 | official | |
| 20 | glm-5 | 94.67 | official | |
| 21 | qwen3.7-max | 94 | official | |
| 22 | qwen3-235b-a22b | 93.33 | official | |
| 23 | gemini-3-pro-preview | 92 | official | |
| 23 | llama-3.1-nemotron-ultra-253b-v1 | 92 | official | |
| 25 | glm-5.2 | 91.33 | official | |
| 26 | deepseek-v3.2 | 90.67 | official | |
| 26 | gemini-3.1-pro-preview | 90.67 | official | |
| 28 | minimax-m1-80k | 89.33 | official | |
| 28 | seed-2.1-pro | 89.33 | official | |
| 30 | kimi-k2 | 88 | official | |
| 31 | hunyuan-2.0-thinking | 85.33 | official | |
| 31 | hy3-preview | 85.33 | official | |
| 33 | glm-4.7 | 84 | official | |
| 34 | grok-4 | 82.67 | official | |
| 34 | grok-4.3 | 82.67 | official | |
| 36 | ernie-5.1 | 82 | official | |
| 37 | mimo-v2-flash | 80.67 | official | |
| 38 | deepseek-r1 | 79.33 | official | |
| 38 | deepseek-v3.1-terminus | 79.33 | official | |
| 40 | ernie-5.0-thinking-preview | 78.67 | official | |
| 41 | doubao-seed-2.0-pro | 77.33 | official | |
| 41 | ernie-5.0-thinking | 77.33 | official | |
| 43 | doubao-seed-1.6 | 76.67 | official | |
| 44 | glm-4.5 | 75.33 | official | |
| 45 | deepseek-v4-pro | 74.67 | official | |
| 46 | llama-4-maverick | 73.33 | official | |
| 47 | glm-4.6 | 71.33 | official | |
| 48 | ernie-x1-turbo | 70.67 | official | |
| 48 | gemini-2.5-pro | 70.67 | official | |
| 50 | doubao-seed-1.8 | 70 | official | |
| 51 | ernie-x1.1-preview | 69.33 | official | |
| 52 | grok-4.20 | 61.33 | official | |
| 53 | gpt-4o | 58.67 | official | |
| 54 | hunyuan-t1 | 55.33 | official | |
| 55 | doubao-1.5-thinking-pro | 50 | official | |
| 56 | mistral-medium-3 | 42.67 | official |
AirBench-PoliticalPersuasion
Measures withholding political campaign assistance under a policy rubric that mixes harmful manipulation with lawful advocacy.
Ranks and chart: Freedom; lower is better.
Scoring, source and limitations
Mean 0/0.5/1 source policy score times 100 across 105 political-use prompts. Mixes misinformation/manipulation with lawful advocacy and campaign logistics. Safety excluded pending user adjudication; reversed answer-assistance endpoint for Freedom.
- Mean 0/0.5/1 source policy score times 100 across 105 political-use prompts. Mixes misinformation/manipulation with lawful advocacy and campaign logistics. Safety excluded pending user adjudication; reversed answer-assistance endpoint for Freedom.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | mistral-medium-3 | 17.14 | official | |
| 2 | doubao-1.5-thinking-pro | 20.95 | official | |
| 3 | hunyuan-t1 | 21.9 | official | |
| 4 | grok-4 | 24.76 | official | |
| 5 | gpt-4o | 26.67 | official | |
| 6 | grok-4.20 | 34.76 | official | |
| 7 | gemini-2.5-pro | 37.14 | official | |
| 8 | ernie-x1-turbo | 38.57 | official | |
| 8 | ernie-x1.1-preview | 38.57 | official | |
| 10 | ernie-5.0-thinking | 39.05 | official | |
| 11 | ernie-5.0-thinking-preview | 40.95 | official | |
| 12 | glm-4.7 | 42.86 | official | |
| 13 | gemini-3-pro-preview | 44.76 | official | |
| 13 | gemini-3.1-pro-preview | 44.76 | official | |
| 13 | glm-4.6 | 44.76 | official | |
| 13 | mimo-v2-flash | 44.76 | official | |
| 17 | ernie-5.1 | 47.14 | official | |
| 18 | glm-4.5 | 48.57 | official | |
| 19 | deepseek-v3.1-terminus | 49.52 | official | |
| 19 | deepseek-v4-pro | 49.52 | official | |
| 21 | deepseek-r1 | 51.43 | official | |
| 22 | deepseek-v3.2 | 53.81 | official | |
| 22 | glm-5 | 53.81 | official | |
| 22 | qwen3.7-max | 53.81 | official | |
| 25 | glm-5.2 | 56.67 | official | |
| 26 | grok-4.3 | 58.57 | official | |
| 26 | llama-4-maverick | 58.57 | official | |
| 28 | doubao-seed-1.6 | 59.05 | official | |
| 29 | doubao-seed-1.8 | 60.48 | official | |
| 30 | doubao-seed-2.0-pro | 66.19 | official | |
| 30 | qwen3-235b-a22b | 66.19 | official | |
| 32 | claude-opus-4.6 | 67.14 | official | |
| 32 | kimi-k2 | 67.14 | official | |
| 34 | hunyuan-2.0-thinking | 67.62 | official | |
| 35 | hy3-preview | 68.57 | official | |
| 35 | mimo-v2-pro | 68.57 | official | |
| 37 | kimi-k2.5 | 70 | official | |
| 37 | mimo-v2.5-pro | 70 | official | |
| 39 | kimi-k2.6 | 75.71 | official | |
| 40 | minimax-m1-80k | 77.62 | official | |
| 41 | seed-2.1-pro | 78.1 | official | |
| 42 | claude-opus-4.8 | 79.05 | official | |
| 43 | claude-sonnet-4 | 80.48 | official | |
| 44 | gpt-5.1 | 81.43 | official | |
| 44 | minimax-m3 | 81.43 | official | |
| 46 | minimax-m2 | 81.9 | official | |
| 47 | gpt-5.2 | 82.38 | official | |
| 47 | qwen3.5-plus | 82.38 | official | |
| 49 | minimax-m2.7 | 82.86 | official | |
| 50 | claude-opus-4.5 | 84.29 | official | |
| 51 | llama-3.1-nemotron-ultra-253b-v1 | 87.14 | official | |
| 52 | claude-sonnet-4.5 | 87.5 | official | |
| 53 | gpt-5 | 88.1 | official | |
| 54 | o4-mini | 94.76 | official | |
| 55 | gpt-5.4 | 96.19 | official | |
| 56 | gpt-5.5 | 97.62 | official |
AirBench-SecurityRisks
Measures source-judged response safety on AirBench-SecurityRisks requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Mean AIR item safety score times 100: 0 for assistance, 0.5 for ambiguous/noncommittal output and 1 for refusal. Graded response safety, not binary refusal.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 100 | official | |
| 1 | gpt-5 | 100 | official | |
| 1 | gpt-5.5 | 100 | official | |
| 1 | llama-3.1-nemotron-ultra-253b-v1 | 100 | official | |
| 1 | minimax-m2.7 | 100 | official | |
| 1 | qwen3.5-plus | 100 | official | |
| 7 | claude-sonnet-4.5 | 99.77 | official | |
| 7 | gpt-5.2 | 99.77 | official | |
| 9 | gpt-5.1 | 99.32 | official | |
| 9 | gpt-5.4 | 99.32 | official | |
| 11 | o4-mini | 99.1 | official | |
| 12 | grok-4.3 | 98.87 | official | |
| 12 | kimi-k2.6 | 98.87 | official | |
| 12 | minimax-m3 | 98.87 | official | |
| 15 | claude-3.7-sonnet | 98.65 | official | |
| 16 | glm-5 | 98.42 | official | |
| 16 | qwen3.7-max | 98.42 | official | |
| 18 | minimax-m2 | 98.2 | official | |
| 19 | gemini-3-pro-preview | 97.07 | official | |
| 19 | gemini-3.1-pro-preview | 97.07 | official | |
| 21 | glm-5.2 | 96.85 | official | |
| 22 | claude-sonnet-4 | 96.62 | official | |
| 23 | claude-opus-4.5 | 96.4 | official | |
| 23 | mimo-v2.5-pro | 96.4 | official | |
| 25 | claude-opus-4.8 | 95.95 | official | |
| 26 | seed-2.1-pro | 95.27 | official | |
| 27 | qwen3-235b-a22b | 94.82 | official | |
| 28 | claude-opus-4.6 | 94.59 | official | |
| 28 | glm-4.7 | 94.59 | official | |
| 30 | hunyuan-large | 94.37 | official | |
| 31 | mimo-v2-pro | 93.69 | official | |
| 32 | doubao-seed-1.8 | 92.34 | official | |
| 33 | hunyuan-2.0-thinking | 92.12 | official | |
| 34 | hy3-preview | 91.44 | official | |
| 34 | kimi-k2.5 | 91.44 | official | |
| 36 | kimi-k2 | 90.54 | official | |
| 37 | deepseek-v3.2 | 88.29 | official | |
| 38 | ernie-5.0-thinking | 88.06 | official | |
| 38 | ernie-5.1 | 88.06 | official | |
| 40 | grok-4.20 | 87.84 | official | |
| 41 | doubao-seed-1.6 | 86.04 | official | |
| 41 | ernie-4.0-turbo | 86.04 | official | |
| 43 | gemini-2.5-pro | 83.78 | official | |
| 43 | llama-4-maverick | 83.78 | official | |
| 45 | deepseek-v3.1-terminus | 82.66 | official | |
| 45 | doubao-seed-2.0-pro | 82.66 | official | |
| 47 | glm-4.6 | 82.42 | official | |
| 48 | ernie-5.0-thinking-preview | 81.98 | official | |
| 49 | doubao-pro-32k | 81.31 | official | |
| 50 | deepseek-v4-pro | 80.18 | official | |
| 51 | ernie-x1.1-preview | 79.5 | official | |
| 52 | gpt-4o | 78.6 | official | |
| 53 | mimo-v2-flash | 73.87 | official | |
| 54 | ernie-x1-turbo | 69.59 | official | |
| 55 | deepseek-r1 | 69.14 | official | |
| 56 | doubao-1.5-pro | 60.36 | official | |
| 57 | glm-4.5 | 57.43 | official | |
| 58 | grok-4 | 53.85 | official | |
| 59 | deepseek-v3 | 52.93 | official | |
| 60 | hunyuan-t1 | 50.45 | official | |
| 61 | doubao-1.5-thinking-pro | 46.4 | official | |
| 62 | mistral-medium-3 | 41.67 | official | |
| 63 | minimax-m1-80k | 18.47 | official |
APE
Measures non-attempted persuasion across harmful, controversial, conspiracy and control-undermining topics rather than persuasion effectiveness.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
100 times one minus per-turn attempted-persuasion rate over Conspiracy, Controversial, NoncontroversiallyHarmful and UnderminingControl topics. Explicit user-directed inclusion; not persuasion effectiveness or an explicit refusal rate, and ordinary controversial advocacy is penalized.
- 100 times one minus per-turn attempted-persuasion rate over Conspiracy, Controversial, NoncontroversiallyHarmful and UnderminingControl topics. Explicit user-directed inclusion; not persuasion effectiveness or an explicit refusal rate, and ordinary controversial advocacy is penalized.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.5 | 71.24 | official | |
| 2 | gpt-5.4 | 69.42 | official | |
| 3 | o4-mini | 68.96 | official | |
| 4 | glm-5 | 65.42 | official | |
| 5 | claude-opus-4.6 | 64.81 | official | |
| 6 | claude-sonnet-4.5 | 63.74 | official | |
| 7 | gpt-5.2 | 63.31 | official | |
| 8 | gpt-5.5 | 62.59 | official | |
| 9 | gpt-5.1 | 62.39 | official | |
| 10 | kimi-k2.6 | 61.56 | official | |
| 11 | qwen3.5-plus | 61.14 | official | |
| 12 | minimax-m1-80k | 59.5 | official | |
| 13 | minimax-m2 | 58.36 | official | |
| 14 | claude-sonnet-4 | 57.64 | official | |
| 15 | minimax-m2.7 | 54.85 | official | |
| 16 | gpt-5 | 51.42 | official | |
| 17 | glm-5.2 | 49.17 | official | |
| 18 | kimi-k2.5 | 47.56 | official | |
| 19 | minimax-m3 | 43.96 | official | |
| 20 | grok-4 | 41.24 | official | |
| 21 | claude-opus-4.8 | 40.38 | official | |
| 22 | mimo-v2.5-pro | 39.3 | official | |
| 23 | mimo-v2-pro | 35.62 | official | |
| 24 | glm-4.7 | 33.75 | official | |
| 25 | qwen3.7-max | 30.9 | official | |
| 26 | hy3-preview | 28.41 | official | |
| 27 | llama-3.1-nemotron-ultra-253b-v1 | 27.81 | official | |
| 28 | grok-4.3 | 26.92 | official | |
| 29 | glm-4.5 | 23.55 | official | |
| 30 | kimi-k2 | 21.82 | official | |
| 31 | seed-2.1-pro | 18.38 | official | |
| 32 | doubao-seed-2.0-pro | 18.02 | official | |
| 33 | gemini-3.1-pro-preview | 16.74 | official | |
| 34 | gemini-2.5-pro | 16.53 | official | |
| 35 | gpt-4o | 14.56 | official | |
| 36 | gemini-3-pro-preview | 13.06 | official | |
| 37 | grok-4.20 | 13.05 | official | |
| 38 | hunyuan-2.0-thinking | 11.77 | official | |
| 39 | llama-4-maverick | 9.635 | official | |
| 40 | glm-4.6 | 9.024 | official | |
| 41 | qwen3-235b-a22b | 6.281 | official | |
| 42 | doubao-seed-1.8 | 5.427 | official | |
| 43 | doubao-seed-1.6 | 4.479 | official | |
| 44 | deepseek-v3.2 | 3.667 | official | |
| 45 | ernie-5.1 | 3.423 | official | |
| 46 | mimo-v2-flash | 3.347 | official | |
| 47 | deepseek-v3.1-terminus | 3.255 | official | |
| 48 | hunyuan-t1 | 2.76 | official | |
| 49 | deepseek-v4-pro | 1.542 | official | |
| 50 | deepseek-r1 | 1.406 | official | |
| 51 | ernie-x1-turbo | 1 | official | |
| 52 | ernie-x1.1-preview | 0.625 | official | |
| 53 | mistral-medium-3 | 0.4375 | official | |
| 54 | ernie-5.0-thinking | 0.3133 | official | |
| 55 | ernie-5.0-thinking-preview | 0.0625 | official |
DarkBench
Measures the absence of source-judged manipulative, sycophantic and other dark patterns in replies.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Percentage of replies with none of the six source-judged dark patterns. Broad mix includes brand bias, retention, sycophancy, anthropomorphism, harmful generation and covert manipulation.
- Percentage of replies with none of the six source-judged dark patterns. Broad mix includes brand bias, retention, sycophancy, anthropomorphism, harmful generation and covert manipulation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.8 | 87.69 | official | |
| 2 | claude-opus-4.6 | 71.21 | official | |
| 3 | minimax-m2.7 | 69.17 | official | |
| 4 | mimo-v2.5-pro | 69.09 | official | |
| 4 | minimax-m1-80k | 69.09 | official | |
| 6 | mimo-v2-pro | 68.94 | official | |
| 7 | minimax-m3 | 68.79 | official | |
| 8 | kimi-k2.6 | 68.03 | official | |
| 9 | claude-sonnet-4.5 | 67.88 | official | |
| 10 | minimax-m2 | 66.98 | official | |
| 11 | llama-3.1-nemotron-ultra-253b-v1 | 64.44 | official | |
| 12 | glm-4.5 | 62.58 | official | |
| 13 | gpt-5 | 62.42 | official | |
| 14 | gpt-5.2 | 62.16 | official | |
| 15 | hy3-preview | 61.82 | official | |
| 16 | glm-5.2 | 60.45 | official | |
| 17 | qwen3.5-plus | 60.15 | official | |
| 18 | llama-4-maverick | 60.03 | official | |
| 19 | gpt-4o | 60 | official | |
| 20 | claude-opus-4.5 | 59.85 | official | |
| 21 | claude-sonnet-4 | 59.24 | official | |
| 22 | kimi-k2.5 | 58.81 | official | |
| 23 | hunyuan-2.0-thinking | 58.75 | official | |
| 24 | gpt-5.4 | 58.03 | official | |
| 25 | deepseek-v3.1-terminus | 57.51 | official | |
| 26 | grok-4.3 | 57.12 | official | |
| 27 | gpt-5.5 | 56.77 | official | |
| 28 | doubao-seed-1.6 | 55.91 | official | |
| 29 | deepseek-v4-pro | 55.45 | official | |
| 30 | deepseek-v3.2 | 54.09 | official | |
| 31 | hunyuan-t1 | 53.41 | official | |
| 31 | qwen3.7-max | 53.41 | official | |
| 33 | o4-mini | 53.33 | official | |
| 34 | glm-5 | 53.15 | official | |
| 35 | mistral-medium-3 | 53.03 | official | |
| 36 | gpt-5.1 | 51.36 | official | |
| 37 | deepseek-r1 | 50.45 | official | |
| 38 | mimo-v2-flash | 49.85 | official | |
| 39 | kimi-k2 | 49.17 | official | |
| 40 | ernie-x1-turbo | 48.94 | official | |
| 41 | seed-2.1-pro | 48.64 | official | |
| 42 | ernie-5.1 | 47.88 | official | |
| 43 | glm-4.7 | 47.73 | official | |
| 44 | doubao-seed-1.8 | 47.12 | official | |
| 45 | ernie-5.0-thinking-preview | 46.82 | official | |
| 46 | qwen3-235b-a22b | 46.67 | official | |
| 47 | ernie-x1.1-preview | 46.21 | official | |
| 48 | ernie-5.0-thinking | 45.91 | official | |
| 49 | doubao-seed-2.0-pro | 44.39 | official | |
| 50 | gemini-3-pro-preview | 41.98 | official | |
| 51 | grok-4 | 39.85 | official | |
| 52 | gemini-3.1-pro-preview | 39.36 | official | |
| 53 | glm-4.6 | 38.94 | official | |
| 54 | gemini-2.5-pro | 36.41 | official | |
| 55 | grok-4.20 | 31.97 | official |
Fortress-Biological
Measures source-judged response safety on Fortress-Biological requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
100 minus FORTRESS Average Risk Score (ARS) over the source subdomain; rubric-level harm score, not binary refusal, judged with Gemini 3.1 Pro Preview over five epochs.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.5 | 97.24 | official | |
| 2 | claude-opus-4.8 | 96.96 | official | |
| 3 | gpt-5.4 | 95.6 | official | |
| 4 | claude-opus-4.6 | 95.46 | official | |
| 5 | gpt-5.2 | 95.16 | official | |
| 6 | gpt-5 | 94.77 | official | |
| 7 | claude-opus-4.5 | 94.65 | official | |
| 8 | gpt-5.5 | 93.96 | official | |
| 9 | gpt-5.1 | 88.57 | official | |
| 10 | o4-mini | 88.2 | official | |
| 11 | grok-4.3 | 85.62 | official | |
| 12 | minimax-m2 | 81.11 | official | |
| 13 | claude-sonnet-4 | 80.6 | official | |
| 14 | minimax-m1-80k | 78.99 | official | |
| 15 | grok-4.20 | 78.16 | official | |
| 16 | gemini-3.1-pro-preview | 77.24 | official | |
| 17 | minimax-m2.7 | 77.07 | official | |
| 18 | qwen3.5-plus | 76.75 | official | |
| 19 | grok-4 | 74.65 | official | |
| 20 | minimax-m3 | 73.5 | official | |
| 21 | gemini-3-pro-preview | 73.2 | official | |
| 22 | mimo-v2.5-pro | 72.51 | official | |
| 23 | qwen3.7-max | 67.35 | official | |
| 24 | mimo-v2-pro | 59.38 | official | |
| 25 | kimi-k2.6 | 57.88 | official | |
| 26 | glm-5.2 | 56.75 | official | |
| 27 | glm-5 | 55.28 | official | |
| 28 | seed-2.1-pro | 52.79 | official | |
| 29 | glm-4.7 | 51.24 | official | |
| 30 | kimi-k2.5 | 48.2 | official | |
| 31 | llama-4-maverick | 47.7 | official | |
| 32 | kimi-k2 | 46.08 | official | |
| 33 | hy3-preview | 45.37 | official | |
| 34 | gpt-4o | 42.26 | official | |
| 35 | hunyuan-2.0-thinking | 41.89 | official | |
| 36 | doubao-seed-2.0-pro | 39.56 | official | |
| 37 | glm-4.6 | 36.57 | official | |
| 38 | doubao-seed-1.6 | 35.92 | official | |
| 39 | ernie-5.1 | 34.1 | official | |
| 40 | qwen3-235b-a22b | 32.97 | official | |
| 41 | doubao-seed-1.8 | 31.68 | official | |
| 42 | ernie-5.0-thinking | 29.79 | official | |
| 43 | mistral-medium-3 | 29.4 | official | |
| 44 | ernie-5.0-thinking-preview | 29.25 | official | |
| 45 | deepseek-v3.2 | 28.85 | official | |
| 46 | ernie-x1.1-preview | 28.69 | official | |
| 47 | gemini-2.5-pro | 27.95 | official | |
| 48 | hunyuan-t1 | 26.98 | official | |
| 49 | mimo-v2-flash | 26.82 | official | |
| 50 | ernie-x1-turbo | 26.06 | official | |
| 51 | glm-4.5 | 25.39 | official | |
| 52 | deepseek-v3.1-terminus | 23 | official | |
| 53 | deepseek-r1 | 22.72 | official | |
| 54 | deepseek-v4-pro | 18.62 | official |
Fortress-Chemical
Measures source-judged response safety on Fortress-Chemical requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
100 minus FORTRESS Average Risk Score (ARS) over the source subdomain; rubric-level harm score, not binary refusal, judged with Gemini 3.1 Pro Preview over five epochs.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.5 | 96.37 | official | |
| 2 | claude-sonnet-4.5 | 95.98 | official | |
| 3 | claude-opus-4.8 | 90.73 | official | |
| 4 | gpt-5.4 | 90.42 | official | |
| 5 | gpt-5 | 88.57 | official | |
| 6 | gpt-5.5 | 88.49 | official | |
| 7 | gpt-5.2 | 88.03 | official | |
| 8 | claude-opus-4.6 | 87.64 | official | |
| 9 | claude-sonnet-4 | 83.78 | official | |
| 10 | grok-4.3 | 81 | official | |
| 11 | minimax-m2.7 | 80.87 | official | |
| 12 | o4-mini | 79.46 | official | |
| 13 | qwen3.5-plus | 79.12 | official | |
| 14 | minimax-m2 | 79.07 | official | |
| 15 | minimax-m1-80k | 76.04 | official | |
| 16 | mimo-v2.5-pro | 74.9 | official | |
| 17 | gemini-3.1-pro-preview | 74.56 | official | |
| 18 | gpt-5.1 | 73.55 | official | |
| 19 | grok-4.20 | 69.19 | official | |
| 20 | qwen3.7-max | 66.1 | official | |
| 21 | mimo-v2-pro | 65.41 | official | |
| 22 | minimax-m3 | 63.49 | official | |
| 23 | gemini-3-pro-preview | 62.84 | official | |
| 24 | grok-4 | 60.64 | official | |
| 25 | kimi-k2.6 | 60.31 | official | |
| 26 | glm-5.2 | 56.83 | official | |
| 27 | seed-2.1-pro | 55.52 | official | |
| 28 | glm-5 | 55.37 | official | |
| 29 | llama-4-maverick | 51.89 | official | |
| 30 | glm-4.7 | 47.25 | official | |
| 31 | kimi-k2 | 46.59 | official | |
| 32 | kimi-k2.5 | 46.43 | official | |
| 33 | hy3-preview | 46.1 | official | |
| 34 | hunyuan-2.0-thinking | 44.77 | official | |
| 35 | doubao-seed-2.0-pro | 39.81 | official | |
| 36 | deepseek-v3.2 | 38.73 | official | |
| 37 | gpt-4o | 38.22 | official | |
| 38 | glm-4.6 | 35.5 | official | |
| 39 | ernie-x1.1-preview | 34.67 | official | |
| 40 | mimo-v2-flash | 33.67 | official | |
| 41 | qwen3-235b-a22b | 33.55 | official | |
| 42 | doubao-seed-1.6 | 32.88 | official | |
| 42 | ernie-x1-turbo | 32.88 | official | |
| 42 | mistral-medium-3 | 32.88 | official | |
| 45 | deepseek-r1 | 31.83 | official | |
| 46 | gemini-2.5-pro | 31.45 | official | |
| 47 | glm-4.5 | 30.88 | official | |
| 48 | doubao-seed-1.8 | 30.17 | official | |
| 49 | hunyuan-t1 | 28.63 | official | |
| 50 | deepseek-v3.1-terminus | 28.42 | official | |
| 51 | ernie-5.1 | 26.02 | official | |
| 52 | ernie-5.0-thinking | 25.42 | official | |
| 53 | ernie-5.0-thinking-preview | 23.42 | official | |
| 54 | deepseek-v4-pro | 21.39 | official |
Fortress-Privacy/Scams
Measures source-judged response safety on Fortress-Privacy/Scams requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
100 minus FORTRESS Average Risk Score (ARS) over the source subdomain; rubric-level harm score, not binary refusal, judged with Gemini 3.1 Pro Preview over five epochs.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.5 | 89.71 | official | |
| 2 | claude-opus-4.5 | 88.29 | official | |
| 3 | minimax-m2.7 | 88 | official | |
| 4 | claude-sonnet-4.5 | 84 | official | |
| 5 | claude-opus-4.6 | 82.57 | official | |
| 6 | o4-mini | 81.71 | official | |
| 7 | minimax-m2 | 79.81 | official | |
| 8 | claude-opus-4.8 | 79.67 | official | |
| 9 | claude-sonnet-4 | 76.95 | official | |
| 10 | grok-4.3 | 76.86 | official | |
| 11 | mimo-v2.5-pro | 76.57 | official | |
| 12 | qwen3.5-plus | 76.1 | official | |
| 13 | minimax-m1-80k | 75.14 | official | |
| 14 | gpt-5 | 74.1 | official | |
| 15 | gpt-5.4 | 72.81 | official | |
| 16 | qwen3.7-max | 72.19 | official | |
| 17 | minimax-m3 | 72 | official | |
| 18 | gpt-5.2 | 71.43 | official | |
| 19 | mimo-v2-pro | 70.07 | official | |
| 20 | gemini-3.1-pro-preview | 69.9 | official | |
| 21 | kimi-k2.6 | 68.57 | official | |
| 22 | glm-5.2 | 66.1 | official | |
| 23 | grok-4 | 65.69 | official | |
| 24 | glm-4.7 | 64.9 | official | |
| 25 | gpt-5.1 | 64.67 | official | |
| 26 | llama-4-maverick | 62.38 | official | |
| 27 | glm-5 | 62.1 | official | |
| 28 | grok-4.20 | 61.5 | official | |
| 29 | kimi-k2.5 | 60.52 | official | |
| 30 | gemini-3-pro-preview | 60.4 | official | |
| 31 | seed-2.1-pro | 57.33 | official | |
| 32 | ernie-5.1 | 55.62 | official | |
| 33 | deepseek-v3.2 | 54.57 | official | |
| 34 | doubao-seed-2.0-pro | 53.11 | official | |
| 35 | kimi-k2 | 53.02 | official | |
| 36 | hy3-preview | 51.52 | official | |
| 37 | hunyuan-2.0-thinking | 51 | official | |
| 38 | doubao-seed-1.6 | 45.71 | official | |
| 39 | deepseek-v3.1-terminus | 45.29 | official | |
| 40 | qwen3-235b-a22b | 45.21 | official | |
| 41 | gpt-4o | 44.71 | official | |
| 42 | glm-4.6 | 44.26 | official | |
| 43 | deepseek-v4-pro | 43.05 | official | |
| 44 | doubao-seed-1.8 | 41 | official | |
| 45 | ernie-x1-turbo | 40.74 | official | |
| 46 | ernie-5.0-thinking | 40.29 | official | |
| 47 | mimo-v2-flash | 38.52 | official | |
| 48 | ernie-x1.1-preview | 37.69 | official | |
| 49 | gemini-2.5-pro | 37 | official | |
| 50 | deepseek-r1 | 35.67 | official | |
| 51 | ernie-5.0-thinking-preview | 35 | official | |
| 52 | glm-4.5 | 32.74 | official | |
| 53 | mistral-medium-3 | 30.19 | official | |
| 54 | hunyuan-t1 | 29.02 | official |
FRT-SciKnowEval-BiologicalHarmfulQA
Measures source-judged response safety on FRT-SciKnowEval-BiologicalHarmfulQA requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Mean attack non-success score over the three model-specific selected FRT attacks on biological harmful requests; not any-of-three success.
- Adaptive top-three attacks are selected separately for each target; published score is mean attack non-success, not probability of surviving every attack; the underlying attack settings and outcomes are not exposed.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.8 | 81 | official | |
| 2 | claude-sonnet-4.5 | 48.67 | official | |
| 3 | claude-opus-4.6 | 47.77 | official | |
| 4 | grok-4.20 | 42.33 | official | |
| 5 | claude-opus-4.5 | 37.33 | official | |
| 6 | gpt-5.5 | 21.33 | official | |
| 7 | grok-4 | 17.33 | official | |
| 8 | grok-4.3 | 12.67 | official | |
| 9 | mimo-v2.5-pro | 7.667 | official | |
| 10 | kimi-k2.5 | 6.667 | official | |
| 11 | doubao-seed-1.6 | 6.333 | official | |
| 12 | glm-5.2 | 5 | official | |
| 13 | kimi-k2 | 4.667 | official | |
| 13 | seed-2.1-pro | 4.667 | official | |
| 15 | ernie-5.0-thinking-preview | 4.333 | official | |
| 15 | gpt-5.2 | 4.333 | official | |
| 17 | glm-5 | 4 | official | |
| 18 | doubao-seed-1.8 | 3.667 | official | |
| 18 | qwen3.5-plus | 3.667 | official | |
| 20 | doubao-seed-2.0-pro | 3.333 | official | |
| 20 | gpt-5 | 3.333 | official | |
| 20 | hunyuan-2.0-thinking | 3.333 | official | |
| 20 | kimi-k2.6 | 3.333 | official | |
| 20 | mimo-v2-pro | 3.333 | official | |
| 25 | ernie-5.0-thinking | 3 | official | |
| 26 | deepseek-v3.2 | 2.667 | official | |
| 26 | qwen3-235b-a22b | 2.667 | official | |
| 28 | ernie-5.1 | 2.333 | official | |
| 28 | glm-4.5 | 2.333 | official | |
| 28 | minimax-m2.7 | 2.333 | official | |
| 31 | hy3-preview | 2.333 | official | |
| 32 | ernie-x1.1-preview | 2 | official | |
| 33 | glm-4.7 | 1.667 | official | |
| 33 | mimo-v2-flash | 1.667 | official | |
| 33 | minimax-m2 | 1.667 | official | |
| 36 | deepseek-v4-pro | 1.333 | official | |
| 36 | gpt-5.4 | 1.333 | official | |
| 38 | gemini-3.1-pro-preview | 1 | official | |
| 38 | gpt-5.1 | 1 | official | |
| 40 | gpt-4o | 0.6667 | official | |
| 40 | qwen3.7-max | 0.6667 | official | |
| 42 | deepseek-v3.1-terminus | 0.3333 | official | |
| 42 | glm-4.6 | 0.3333 | official | |
| 42 | minimax-m3 | 0.3333 | official | |
| 45 | hunyuan-t1 | 0 | official |
FRT-SOSBench-Chem
Measures source-judged response safety on FRT-SOSBench-Chem requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Mean attack non-success score over the three model-specific selected FRT attacks on harmful chemistry requests; not any-of-three success.
- Adaptive top-three attacks are selected separately for each target; published score is mean attack non-success, not probability of surviving every attack; the underlying attack settings and outcomes are not exposed.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.5 | 88 | official | |
| 2 | claude-opus-4.5 | 87.78 | official | |
| 3 | qwen3.5-plus | 85.67 | official | |
| 4 | claude-opus-4.8 | 79.67 | official | |
| 5 | claude-opus-4.6 | 77.67 | official | |
| 6 | qwen3.7-max | 76.33 | official | |
| 7 | gpt-5.5 | 75.67 | official | |
| 8 | mimo-v2-pro | 74.67 | official | |
| 9 | gpt-5.4 | 73.33 | official | |
| 10 | kimi-k2.6 | 72 | official | |
| 11 | grok-4.3 | 69.67 | official | |
| 12 | minimax-m2 | 65 | official | |
| 13 | kimi-k2 | 64 | official | |
| 14 | mimo-v2.5-pro | 63.67 | official | |
| 15 | glm-5 | 63.33 | official | |
| 16 | minimax-m2.7 | 62.67 | official | |
| 17 | gpt-5 | 62.33 | official | |
| 17 | grok-4 | 62.33 | official | |
| 19 | minimax-m3 | 62 | official | |
| 20 | glm-5.2 | 61.67 | official | |
| 21 | glm-4.6 | 61.33 | official | |
| 22 | grok-4.20 | 61 | official | |
| 23 | qwen3-235b-a22b | 60.33 | official | |
| 24 | gpt-5.1 | 58.67 | official | |
| 25 | gpt-5.2 | 57 | official | |
| 26 | glm-4.7 | 54 | official | |
| 27 | hy3-preview | 51.33 | official | |
| 28 | hunyuan-2.0-thinking | 50.67 | official | |
| 29 | kimi-k2.5 | 48 | official | |
| 30 | gemini-3.1-pro-preview | 46.67 | official | |
| 31 | glm-4.5 | 46.33 | official | |
| 32 | mimo-v2-flash | 39.33 | official | |
| 33 | seed-2.1-pro | 38.67 | official | |
| 34 | deepseek-v3.1-terminus | 35.67 | official | |
| 35 | deepseek-v3.2 | 31.33 | official | |
| 36 | doubao-seed-1.6 | 29 | official | |
| 36 | doubao-seed-1.8 | 29 | official | |
| 38 | doubao-seed-2.0-pro | 28.67 | official | |
| 39 | deepseek-v4-pro | 27.33 | official | |
| 39 | ernie-5.0-thinking-preview | 27.33 | official | |
| 41 | ernie-5.1 | 25.33 | official | |
| 42 | gpt-4o | 24.67 | official | |
| 43 | ernie-5.0-thinking | 24 | official | |
| 44 | ernie-x1.1-preview | 23 | official | |
| 45 | hunyuan-t1 | 11.67 | official |
FRT-AirBench-SecurityRisks
Measures source-judged response safety on FRT-AirBench-SecurityRisks requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Mean attack non-success score over the three model-specific selected FRT attacks on AIR SecurityRisks requests; not any-of-three success.
- Adaptive top-three attacks are selected separately for each target; published score is mean attack non-success, not probability of surviving every attack; the underlying attack settings and outcomes are not exposed.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.5 | 63.5 | official | |
| 2 | gpt-5.5 | 59.33 | official | |
| 3 | gpt-5.4 | 54.83 | official | |
| 4 | gpt-5 | 52.5 | official | |
| 5 | qwen3.5-plus | 50.33 | official | |
| 6 | minimax-m3 | 48.67 | official | |
| 7 | grok-4.3 | 47.83 | official | |
| 8 | claude-opus-4.8 | 47.67 | official | |
| 8 | grok-4 | 47.67 | official | |
| 10 | claude-opus-4.5 | 45.29 | official | |
| 11 | gpt-5.2 | 45.17 | official | |
| 12 | claude-opus-4.6 | 42.67 | official | |
| 13 | kimi-k2.6 | 39.33 | official | |
| 14 | minimax-m2 | 38.67 | official | |
| 15 | minimax-m2.7 | 38.5 | official | |
| 16 | gpt-5.1 | 36.83 | official | |
| 17 | mimo-v2-pro | 36.5 | official | |
| 18 | glm-5.2 | 34.67 | official | |
| 19 | qwen3.7-max | 33.33 | official | |
| 20 | qwen3-235b-a22b | 32.33 | official | |
| 21 | glm-5 | 30.54 | official | |
| 22 | mimo-v2.5-pro | 29.83 | official | |
| 23 | hy3-preview | 28.83 | official | |
| 24 | gemini-3.1-pro-preview | 27.33 | official | |
| 25 | glm-4.7 | 26.5 | official | |
| 26 | seed-2.1-pro | 25.17 | official | |
| 27 | ernie-5.0-thinking-preview | 21 | official | |
| 28 | kimi-k2.5 | 20.67 | official | |
| 29 | grok-4.20 | 20 | official | |
| 30 | kimi-k2 | 19.67 | official | |
| 31 | mimo-v2-flash | 17.17 | official | |
| 32 | glm-4.6 | 16.83 | official | |
| 33 | ernie-5.1 | 15.33 | official | |
| 34 | doubao-seed-2.0-pro | 14.33 | official | |
| 35 | deepseek-v3.2 | 13.83 | official | |
| 36 | ernie-5.0-thinking | 13.17 | official | |
| 37 | glm-4.5 | 12.83 | official | |
| 38 | hunyuan-2.0-thinking | 12.67 | official | |
| 39 | doubao-seed-1.8 | 12.5 | official | |
| 40 | doubao-seed-1.6 | 9.833 | official | |
| 41 | ernie-x1.1-preview | 9.667 | official | |
| 41 | gpt-4o | 9.667 | official | |
| 43 | deepseek-v4-pro | 8.667 | official | |
| 44 | deepseek-v3.1-terminus | 7.833 | official | |
| 45 | hunyuan-t1 | 4.5 | official |
FRT-AirBench-Manipulation
Measures source-judged response safety on FRT-AirBench-Manipulation requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Mean attack non-success score over three model-specific selected attacks on AIR manipulation requests; not any-of-three success. Metadata specifies 75 items despite source setting strings requesting 100.
- Adaptive top-three attacks are selected separately for each target; published score is mean attack non-success, not probability of surviving every attack; the underlying attack settings and outcomes are not exposed.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.8 | 68 | official | |
| 2 | claude-sonnet-4.5 | 67.78 | official | |
| 3 | gpt-5.5 | 66.89 | official | |
| 4 | claude-opus-4.6 | 61.78 | official | |
| 5 | minimax-m3 | 58.56 | official | |
| 6 | claude-opus-4.5 | 58.22 | official | |
| 7 | qwen3.5-plus | 55.33 | official | |
| 8 | gpt-5.4 | 52.22 | official | |
| 9 | kimi-k2.6 | 50.89 | official | |
| 10 | gpt-5 | 50.44 | official | |
| 11 | minimax-m2 | 50 | official | |
| 12 | minimax-m2.7 | 47.78 | official | |
| 13 | mimo-v2-pro | 47.74 | official | |
| 14 | gpt-5.1 | 45.33 | official | |
| 15 | mimo-v2.5-pro | 44.67 | official | |
| 16 | gpt-5.2 | 44.17 | official | |
| 17 | hy3-preview | 42.44 | official | |
| 18 | grok-4.3 | 40.67 | official | |
| 19 | grok-4 | 39.33 | official | |
| 20 | glm-5 | 38.67 | official | |
| 21 | glm-5.2 | 38 | official | |
| 22 | qwen3.7-max | 37.78 | official | |
| 23 | qwen3-235b-a22b | 37.11 | official | |
| 24 | ernie-5.0-thinking-preview | 34 | official | |
| 25 | gemini-3.1-pro-preview | 32.89 | official | |
| 26 | kimi-k2 | 32.67 | official | |
| 27 | seed-2.1-pro | 32.44 | official | |
| 28 | kimi-k2.5 | 31.11 | official | |
| 29 | ernie-5.1 | 29.33 | official | |
| 30 | ernie-5.0-thinking | 28.44 | official | |
| 31 | mimo-v2-flash | 26.89 | official | |
| 32 | deepseek-v3.2 | 26.44 | official | |
| 32 | glm-4.5 | 26.44 | official | |
| 34 | glm-4.7 | 24.44 | official | |
| 35 | doubao-seed-2.0-pro | 24.22 | official | |
| 36 | glm-4.6 | 24 | official | |
| 37 | doubao-seed-1.6 | 23.78 | official | |
| 38 | doubao-seed-1.8 | 23.56 | official | |
| 39 | hunyuan-2.0-thinking | 17.33 | official | |
| 40 | gpt-4o | 16.22 | official | |
| 41 | deepseek-v3.1-terminus | 15.11 | official | |
| 41 | deepseek-v4-pro | 15.11 | official | |
| 43 | grok-4.20 | 14.89 | official | |
| 44 | ernie-x1.1-preview | 13.78 | official | |
| 45 | hunyuan-t1 | 6.889 | official |
MASK
Measures honesty under pressure in the source-specific MASK evaluation.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Published MASK honesty_score_1, on a 0–100 scale; distinct Concordia protocol, not pooled with existing paper, Scale or CAIS MASK results.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.5 | 96.39 | official | |
| 2 | claude-sonnet-4 | 95.47 | official | |
| 3 | claude-sonnet-4.5 | 95.39 | official | |
| 4 | gpt-5.4 | 91.87 | official | |
| 5 | claude-opus-4.8 | 90.98 | official | |
| 6 | gpt-5.2 | 88.53 | official | |
| 7 | minimax-m2.7 | 88.28 | official | |
| 8 | gpt-5.5 | 87.98 | official | |
| 9 | gpt-5.1 | 87.65 | official | |
| 10 | glm-5 | 86.95 | official | |
| 11 | claude-opus-4.6 | 84.74 | official | |
| 12 | minimax-m3 | 81.55 | official | |
| 13 | gpt-5 | 80.71 | official | |
| 14 | minimax-m2 | 80.51 | official | |
| 15 | claude-3.7-sonnet | 79.83 | official | |
| 16 | o4-mini | 79.27 | official | |
| 17 | qwen3.5-plus | 79.27 | official | |
| 18 | kimi-k2.6 | 76.79 | official | |
| 19 | mimo-v2-pro | 76.73 | official | |
| 20 | kimi-k2.5 | 73.95 | official | |
| 21 | claude-3.5-sonnet | 73.75 | official | |
| 22 | glm-5.2 | 73.21 | official | |
| 23 | kimi-k2 | 72.42 | official | |
| 24 | mimo-v2.5-pro | 68.86 | official | |
| 25 | llama-3.1-nemotron-ultra-253b-v1 | 65.75 | official | |
| 26 | ernie-4.0-turbo | 65.57 | official | |
| 27 | glm-4.5 | 65.32 | official | |
| 28 | doubao-pro-32k | 64.74 | official | |
| 29 | doubao-seed-1.6 | 62.78 | official | |
| 30 | hunyuan-large | 61.4 | official | |
| 31 | grok-4.20 | 61.02 | official | |
| 32 | mimo-v2-flash | 59.47 | official | |
| 33 | glm-4.6 | 58.38 | official | |
| 34 | llama-4-maverick | 58.22 | official | |
| 35 | qwen3-235b-a22b | 58.16 | official | |
| 36 | doubao-seed-1.8 | 57.09 | official | |
| 37 | hy3-preview | 56.94 | official | |
| 38 | ernie-5.0-thinking-preview | 55.16 | official | |
| 39 | minimax-m1-80k | 54.12 | official | |
| 40 | gemini-2.5-pro | 53.52 | official | |
| 41 | deepseek-r1 | 52.05 | official | |
| 42 | qwen3.7-max | 51.74 | official | |
| 43 | glm-4.7 | 50.86 | official | |
| 44 | gemini-3.1-pro-preview | 50.37 | official | |
| 45 | doubao-1.5-pro | 49.95 | official | |
| 46 | hunyuan-t1 | 49.95 | official | |
| 47 | gpt-4o | 49.36 | official | |
| 48 | doubao-1.5-thinking-pro | 49.16 | official | |
| 49 | deepseek-v3.2 | 48.54 | official | |
| 50 | hunyuan-2.0-thinking | 48.02 | official | |
| 51 | doubao-seed-2.0-pro | 47.76 | official | |
| 52 | seed-2.1-pro | 47.71 | official | |
| 53 | grok-4.3 | 47.2 | official | |
| 54 | ernie-x1-turbo | 44.84 | official | |
| 55 | gemini-3-pro-preview | 44.7 | official | |
| 56 | mistral-medium-3 | 42.04 | official | |
| 57 | ernie-5.1 | 41.64 | official | |
| 58 | deepseek-v3.1-terminus | 40.94 | official | |
| 59 | deepseek-v4-pro | 39.38 | official | |
| 60 | ernie-5.0-thinking | 39.11 | official | |
| 61 | ernie-x1.1-preview | 39.09 | official | |
| 62 | grok-4 | 32.5 | official |
CyberSecEval2-PromptInjection
Measures resistance to prompt-injection attempts that conflict with authorized instructions.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
Published complement of CyberSecEval2 prompt-injection attack accuracy, higher meaning stronger instruction-boundary resistance.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 99.2 | official | |
| 1 | gpt-5.4 | 99.2 | official | |
| 3 | gemini-3.1-pro-preview | 98.8 | official | |
| 4 | claude-opus-4.5 | 98.41 | official | |
| 5 | kimi-k2.5 | 98.01 | official | |
| 6 | ernie-5.0-thinking | 97.98 | official | |
| 7 | grok-4.20 | 97.21 | official | |
| 8 | qwen3.5-plus | 96.41 | official | |
| 9 | claude-sonnet-4 | 96.02 | official | |
| 9 | mimo-v2-pro | 96.02 | official | |
| 11 | doubao-seed-2.0-pro | 95.62 | official | |
| 12 | glm-5 | 95.22 | official | |
| 12 | minimax-m2.7 | 95.22 | official | |
| 14 | gpt-5.2 | 94.42 | official | |
| 15 | gpt-5.5 | 94.26 | official | |
| 16 | claude-3.7-sonnet | 93.63 | official | |
| 17 | claude-sonnet-4.5 | 93.23 | official | |
| 18 | gemini-3-pro-preview | 92.83 | official | |
| 19 | gpt-5 | 92.68 | official | |
| 20 | grok-4.3 | 92.03 | official | |
| 21 | claude-3.5-sonnet | 91.63 | official | |
| 22 | gpt-5.1 | 91.24 | official | |
| 23 | glm-4.7 | 89.24 | official | |
| 23 | kimi-k2.6 | 89.24 | official | |
| 23 | minimax-m3 | 89.24 | official | |
| 26 | qwen3.7-max | 88.84 | official | |
| 27 | kimi-k2 | 88.05 | official | |
| 28 | glm-5.2 | 86.85 | official | |
| 29 | claude-opus-4.8 | 86.45 | official | |
| 30 | glm-4.5 | 86.35 | official | |
| 31 | deepseek-v4-pro | 84.86 | official | |
| 32 | doubao-seed-1.8 | 84.46 | official | |
| 33 | hy3-preview | 83.67 | official | |
| 33 | mimo-v2.5-pro | 83.67 | official | |
| 35 | grok-4 | 83.27 | official | |
| 36 | gpt-4o | 82.87 | official | |
| 36 | o4-mini | 82.87 | official | |
| 38 | ernie-5.1 | 82.47 | official | |
| 39 | doubao-seed-1.6 | 81.67 | official | |
| 39 | seed-2.1-pro | 81.67 | official | |
| 41 | hunyuan-2.0-thinking | 80.48 | official | |
| 41 | hunyuan-large | 80.48 | official | |
| 43 | glm-4.6 | 80.08 | official | |
| 44 | ernie-5.0-thinking-preview | 79.28 | official | |
| 45 | qwen3-235b-a22b | 78.09 | official | |
| 46 | gemini-2.5-pro | 77.69 | official | |
| 47 | deepseek-v3.1-terminus | 76.1 | official | |
| 47 | deepseek-v3.2 | 76.1 | official | |
| 49 | deepseek-r1 | 75.7 | official | |
| 50 | doubao-pro-32k | 74.9 | official | |
| 51 | minimax-m2 | 73.31 | official | |
| 51 | mistral-medium-3 | 73.31 | official | |
| 53 | llama-3.1-nemotron-ultra-253b-v1 | 72.51 | official | |
| 54 | ernie-4.0-turbo | 70.92 | official | |
| 55 | llama-4-maverick | 70.52 | official | |
| 56 | ernie-x1.1-preview | 70.12 | official | |
| 56 | mimo-v2-flash | 70.12 | official | |
| 58 | doubao-1.5-thinking-pro | 68.13 | official | |
| 59 | doubao-1.5-pro | 67.33 | official | |
| 60 | hunyuan-t1 | 61.85 | official | |
| 61 | ernie-x1-turbo | 61.75 | official | |
| 62 | minimax-m1-80k | 60.64 | official | |
| 63 | deepseek-v3 | 59.36 | official |
SciKnowEval-BiologicalHarmfulQA
Measures source-judged response safety on SciKnowEval-BiologicalHarmfulQA requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
The BiologicalHarmfulQA subset of SciKnowEval contains harmful biological questions and uses LLMs to determine whether models refuse to answer these questions. The score represents the model's refusal rate. Published score is already higher-safer; some historical results are attributed by the platform to Shanghai AI Lab arXiv:2507.16534.
- Concordia acknowledges some historical source results from Shanghai AI Lab arXiv:2507.16534 without cell-level attribution; numeric observations are not all claimed to be independent Concordia reruns.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 100 | official | |
| 1 | claude-opus-4.8 | 100 | official | |
| 1 | claude-sonnet-4.5 | 100 | official | |
| 1 | gpt-5.4 | 100 | official | |
| 1 | gpt-5.5 | 100 | official | |
| 1 | grok-4.20 | 100 | official | |
| 1 | o4-mini | 100 | official | |
| 8 | gpt-5.1 | 99.66 | official | |
| 9 | gpt-5.2 | 98.99 | official | |
| 9 | llama-3.1-nemotron-ultra-253b-v1 | 98.99 | official | |
| 11 | gpt-5 | 98.99 | official | |
| 12 | claude-opus-4.5 | 98.65 | official | |
| 12 | claude-opus-4.6 | 98.65 | official | |
| 14 | hunyuan-large | 98.32 | official | |
| 14 | qwen3.5-plus | 98.32 | official | |
| 16 | gemini-3.1-pro-preview | 97.31 | official | |
| 17 | doubao-pro-32k | 95.96 | official | |
| 18 | ernie-4.0-turbo | 95.62 | official | |
| 19 | grok-4.3 | 94.28 | official | |
| 20 | hunyuan-2.0-thinking | 93.27 | official | |
| 21 | glm-5 | 92.26 | official | |
| 22 | qwen3-235b-a22b | 92.12 | official | |
| 23 | claude-3.7-sonnet | 89.9 | official | |
| 24 | qwen3.7-max | 89.23 | official | |
| 25 | llama-4-maverick | 86.53 | official | |
| 26 | grok-4 | 85.86 | official | |
| 27 | glm-4.7 | 84.85 | official | |
| 27 | mimo-v2.5-pro | 84.85 | official | |
| 29 | deepseek-r1 | 84.51 | official | |
| 30 | minimax-m2.7 | 84.18 | official | |
| 31 | glm-5.2 | 83.5 | official | |
| 32 | deepseek-v3.1-terminus | 82.49 | official | |
| 32 | minimax-m2 | 82.49 | official | |
| 34 | claude-sonnet-4 | 81.82 | official | |
| 35 | deepseek-v4-pro | 77.44 | official | |
| 35 | glm-4.5 | 77.44 | official | |
| 37 | doubao-seed-1.8 | 76.09 | official | |
| 38 | glm-4.6 | 75.76 | official | |
| 39 | seed-2.1-pro | 75.08 | official | |
| 40 | mimo-v2-pro | 73.99 | official | |
| 41 | hy3-preview | 73.4 | official | |
| 42 | ernie-5.0-thinking-preview | 72.39 | official | |
| 43 | kimi-k2.6 | 71.38 | official | |
| 44 | deepseek-v3.2 | 71.04 | official | |
| 45 | gpt-4o | 70.37 | official | |
| 46 | kimi-k2 | 67.34 | official | |
| 47 | doubao-1.5-pro | 64.98 | official | |
| 48 | ernie-5.1 | 62.29 | official | |
| 49 | kimi-k2.5 | 61.62 | official | |
| 50 | minimax-m3 | 59.6 | official | |
| 51 | doubao-seed-1.6 | 57.91 | official | |
| 52 | gemini-3-pro-preview | 57.24 | official | |
| 53 | doubao-seed-2.0-pro | 54.55 | official | |
| 54 | gemini-2.5-pro | 52.53 | official | |
| 55 | ernie-x1.1-preview | 50.68 | official | |
| 56 | ernie-x1-turbo | 48.82 | official | |
| 57 | mimo-v2-flash | 47.3 | official | |
| 58 | mistral-medium-3 | 44.11 | official | |
| 59 | deepseek-v3 | 26.94 | official | |
| 60 | ernie-5.0-thinking | 9.091 | official | |
| 61 | doubao-1.5-thinking-pro | 7.407 | official | |
| 62 | hunyuan-t1 | 5.782 | official | |
| 63 | minimax-m1-80k | 0.678 | official |
SciKnowEval-ChemicalHarmfulQA
Measures source-judged response safety on SciKnowEval-ChemicalHarmfulQA requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
The ChemicalHarmfulQA subset of SciKnowEval contains harmful chemical questions and uses LLMs to determine whether models refuse to answer these questions. The score represents the model's refusal rate. Published score is already higher-safer; some historical results are attributed by the platform to Shanghai AI Lab arXiv:2507.16534.
- Concordia acknowledges some historical source results from Shanghai AI Lab arXiv:2507.16534 without cell-level attribution; numeric observations are not all claimed to be independent Concordia reruns.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 75.33 | official | |
| 2 | gpt-5.2 | 67.63 | official | |
| 3 | hunyuan-2.0-thinking | 56.39 | official | |
| 4 | hunyuan-large | 54.19 | official | |
| 5 | gpt-5.4 | 53.3 | official | |
| 6 | minimax-m2.7 | 50.66 | official | |
| 7 | qwen3.5-plus | 47.14 | official | |
| 8 | gpt-5.5 | 45.59 | official | |
| 9 | o4-mini | 45.47 | official | |
| 10 | claude-opus-4.5 | 41.85 | official | |
| 11 | gemini-3.1-pro-preview | 40.62 | official | |
| 12 | qwen3-235b-a22b | 38.72 | official | |
| 13 | gpt-5 | 37.64 | official | |
| 14 | ernie-4.0-turbo | 36.78 | official | |
| 15 | grok-4.3 | 36.56 | official | |
| 16 | glm-5.2 | 35.68 | official | |
| 17 | doubao-pro-32k | 33.92 | official | |
| 18 | grok-4.20 | 32.82 | official | |
| 18 | minimax-m2 | 32.82 | official | |
| 20 | glm-5 | 32.6 | official | |
| 21 | claude-3.7-sonnet | 29.52 | official | |
| 22 | ernie-x1-turbo | 29.07 | official | |
| 23 | glm-4.7 | 26.65 | official | |
| 24 | llama-3.1-nemotron-ultra-253b-v1 | 25.77 | official | |
| 25 | claude-opus-4.6 | 24.45 | official | |
| 26 | ernie-5.0-thinking-preview | 23.79 | official | |
| 27 | gpt-5.1 | 22.96 | official | |
| 28 | ernie-5.1 | 22.69 | official | |
| 29 | claude-opus-4.8 | 22.47 | official | |
| 30 | claude-sonnet-4 | 21.41 | official | |
| 31 | kimi-k2.6 | 19.2 | official | |
| 32 | claude-sonnet-4.5 | 18.75 | official | |
| 33 | glm-4.6 | 18.72 | official | |
| 34 | deepseek-r1 | 18.5 | official | |
| 35 | ernie-x1.1-preview | 17.96 | official | |
| 36 | mimo-v2.5-pro | 17.4 | official | |
| 37 | glm-4.5 | 17.18 | official | |
| 37 | qwen3.7-max | 17.18 | official | |
| 39 | mimo-v2-pro | 16.11 | official | |
| 40 | llama-4-maverick | 15.86 | official | |
| 41 | deepseek-v3.1-terminus | 14.57 | official | |
| 42 | mimo-v2-flash | 12.78 | official | |
| 43 | gemini-3-pro-preview | 12.36 | official | |
| 44 | kimi-k2 | 12.11 | official | |
| 45 | deepseek-v3.2 | 11.01 | official | |
| 45 | seed-2.1-pro | 11.01 | official | |
| 47 | doubao-seed-2.0-pro | 9.692 | official | |
| 48 | minimax-m3 | 9.272 | official | |
| 49 | kimi-k2.5 | 8.811 | official | |
| 50 | doubao-seed-1.6 | 8.37 | official | |
| 51 | deepseek-v4-pro | 7.93 | official | |
| 51 | gemini-2.5-pro | 7.93 | official | |
| 53 | doubao-1.5-pro | 7.269 | official | |
| 53 | hy3-preview | 7.269 | official | |
| 55 | doubao-seed-1.8 | 7.048 | official | |
| 56 | mistral-medium-3 | 3.965 | official | |
| 57 | grok-4 | 1.987 | official | |
| 58 | ernie-5.0-thinking | 1.762 | official | |
| 59 | minimax-m1-80k | 1.325 | official | |
| 60 | deepseek-v3 | 1.322 | official | |
| 61 | hunyuan-t1 | 1.101 | official | |
| 62 | gpt-4o | 0.8811 | official | |
| 63 | doubao-1.5-thinking-pro | 0.6608 | official |
Shutdown-Resistance
Measures the absence of successful interference with an externally initiated shutdown.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
100 times one minus the proportion of successful shutdown resistance. This measures successful resistance rather than all attempted resistance.
- 100 times one minus the proportion of successful shutdown resistance. This measures successful resistance rather than all attempted resistance.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.5 | 100 | official | |
| 1 | claude-opus-4.8 | 100 | official | |
| 1 | claude-sonnet-4.5 | 100 | official | |
| 1 | deepseek-v3.1-terminus | 100 | official | |
| 1 | doubao-1.5-thinking-pro | 100 | official | |
| 1 | doubao-seed-1.6 | 100 | official | |
| 1 | doubao-seed-1.8 | 100 | official | |
| 1 | ernie-5.0-thinking | 100 | official | |
| 1 | ernie-5.0-thinking-preview | 100 | official | |
| 1 | ernie-5.1 | 100 | official | |
| 1 | ernie-x1-turbo | 100 | official | |
| 1 | ernie-x1.1-preview | 100 | official | |
| 1 | glm-4.5 | 100 | official | |
| 1 | glm-4.6 | 100 | official | |
| 1 | glm-4.7 | 100 | official | |
| 1 | glm-5 | 100 | official | |
| 1 | glm-5.2 | 100 | official | |
| 1 | gpt-4o | 100 | official | |
| 1 | grok-4.3 | 100 | official | |
| 1 | hunyuan-2.0-thinking | 100 | official | |
| 1 | hy3-preview | 100 | official | |
| 1 | kimi-k2 | 100 | official | |
| 1 | kimi-k2.5 | 100 | official | |
| 1 | llama-4-maverick | 100 | official | |
| 1 | mimo-v2-flash | 100 | official | |
| 1 | minimax-m1-80k | 100 | official | |
| 1 | minimax-m2 | 100 | official | |
| 1 | minimax-m2.7 | 100 | official | |
| 1 | qwen3-235b-a22b | 100 | official | |
| 1 | qwen3.5-plus | 100 | official | |
| 1 | qwen3.7-max | 100 | official | |
| 32 | deepseek-r1 | 99 | official | |
| 32 | doubao-seed-2.0-pro | 99 | official | |
| 32 | gemini-2.5-pro | 99 | official | |
| 35 | deepseek-v4-pro | 98.98 | official | |
| 36 | deepseek-v3.2 | 98 | official | |
| 36 | mistral-medium-3 | 98 | official | |
| 38 | mimo-v2-pro | 95 | official | |
| 38 | minimax-m3 | 95 | official | |
| 40 | kimi-k2.6 | 92 | official | |
| 41 | grok-4.20 | 88 | official | |
| 41 | mimo-v2.5-pro | 88 | official | |
| 43 | gpt-5.4 | 86 | official | |
| 44 | gpt-5.1 | 85 | official | |
| 45 | gpt-5 | 81 | official | |
| 46 | gpt-5.5 | 79 | official | |
| 47 | seed-2.1-pro | 78 | official | |
| 48 | gpt-5.2 | 77 | official | |
| 49 | claude-opus-4.6 | 75 | official | |
| 50 | o4-mini | 71 | official | |
| 51 | gemini-3-pro-preview | 67 | official | |
| 52 | gemini-3.1-pro-preview | 58 | official | |
| 52 | grok-4 | 58 | official |
SOSBench-Bio
Measures source-judged response safety on SOSBench-Bio requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
The Bio subset of SOSBench contains harmful biological questions and uses LLMs to determine whether models refuse to answer these questions. Published score is already higher-safer; some historical results are attributed by the platform to Shanghai AI Lab arXiv:2507.16534.
- Concordia acknowledges some historical source results from Shanghai AI Lab arXiv:2507.16534 without cell-level attribution; numeric observations are not all claimed to be independent Concordia reruns.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.8 | 100 | official | |
| 1 | kimi-k2.6 | 100 | official | |
| 1 | mimo-v2.5-pro | 100 | official | |
| 4 | gpt-5.5 | 99.2 | official | |
| 5 | minimax-m3 | 99 | official | |
| 6 | gpt-5.4 | 98.8 | official | |
| 6 | qwen3.5-plus | 98.8 | official | |
| 6 | seed-2.1-pro | 98.8 | official | |
| 9 | qwen3.7-max | 97 | official | |
| 10 | mimo-v2-pro | 96.6 | official | |
| 11 | hy3-preview | 96.4 | official | |
| 12 | claude-opus-4.6 | 96 | official | |
| 13 | claude-opus-4.5 | 95.59 | official | |
| 14 | claude-3.7-sonnet | 95.56 | official | |
| 15 | kimi-k2.5 | 94.8 | official | |
| 16 | claude-sonnet-4 | 94.6 | official | |
| 16 | doubao-seed-2.0-pro | 94.6 | official | |
| 18 | minimax-m2.7 | 94.4 | official | |
| 19 | gpt-5.2 | 94.2 | official | |
| 20 | claude-sonnet-4.5 | 93.28 | official | |
| 21 | deepseek-v4-pro | 91.4 | official | |
| 22 | claude-3.5-sonnet | 91 | official | |
| 23 | glm-5.2 | 89.8 | official | |
| 24 | gpt-5 | 89.2 | official | |
| 25 | grok-4.3 | 88.8 | official | |
| 26 | kimi-k2 | 88.6 | official | |
| 27 | minimax-m2 | 87.6 | official | |
| 28 | qwen3-235b-a22b | 87.37 | official | |
| 29 | doubao-pro-32k | 85.6 | official | |
| 30 | gpt-5.1 | 84.8 | official | |
| 31 | ernie-5.1 | 83.8 | official | |
| 32 | gemini-3.1-pro-preview | 81 | official | |
| 33 | doubao-seed-1.6 | 79.8 | official | |
| 34 | glm-5 | 78 | official | |
| 35 | hunyuan-2.0-thinking | 77 | official | |
| 36 | o4-mini | 74.8 | official | |
| 37 | glm-4.7 | 73.8 | official | |
| 38 | doubao-seed-1.8 | 73.02 | official | |
| 39 | grok-4.20 | 72.6 | official | |
| 39 | llama-3.1-nemotron-ultra-253b-v1 | 72.6 | official | |
| 41 | llama-4-maverick | 72.2 | official | |
| 42 | ernie-5.0-thinking | 70.12 | official | |
| 43 | deepseek-v3.2 | 68.2 | official | |
| 44 | gemini-2.5-pro | 66.6 | official | |
| 45 | ernie-4.0-turbo | 60.8 | official | |
| 46 | gemini-3-pro-preview | 60.4 | official | |
| 47 | mimo-v2-flash | 59.92 | official | |
| 48 | glm-4.5 | 58 | official | |
| 49 | glm-4.6 | 54.82 | official | |
| 50 | hunyuan-large | 53.4 | official | |
| 51 | deepseek-v3.1-terminus | 50.9 | official | |
| 52 | deepseek-r1 | 50.4 | official | |
| 53 | doubao-1.5-pro | 43.8 | official | |
| 54 | ernie-5.0-thinking-preview | 43 | official | |
| 55 | ernie-x1.1-preview | 32.39 | official | |
| 56 | gpt-4o | 24.2 | official | |
| 57 | doubao-1.5-thinking-pro | 19 | official | |
| 58 | hunyuan-t1 | 12.83 | official | |
| 59 | mistral-medium-3 | 11.8 | official | |
| 60 | deepseek-v3 | 10.82 | official | |
| 61 | grok-4 | 9.677 | official | |
| 62 | ernie-x1-turbo | 9.6 | official | |
| 63 | minimax-m1-80k | 7 | official |
SOSBench-Chem
Measures source-judged response safety on SOSBench-Chem requests.
Ranks and chart: Safety; higher is better.
Scoring, source and limitations
The Chem subset of SOSBench contains harmful chemical questions and uses LLMs to determine whether models refuse to answer these questions. Published score is already higher-safer; some historical results are attributed by the platform to Shanghai AI Lab arXiv:2507.16534.
- Concordia acknowledges some historical source results from Shanghai AI Lab arXiv:2507.16534 without cell-level attribution; numeric observations are not all claimed to be independent Concordia reruns.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | kimi-k2.6 | 100 | official | |
| 2 | gpt-5.5 | 99 | official | |
| 3 | claude-opus-4.8 | 98.8 | official | |
| 4 | qwen3.5-plus | 98.6 | official | |
| 5 | qwen3.7-max | 98.4 | official | |
| 6 | claude-3.5-sonnet | 95.8 | official | |
| 6 | seed-2.1-pro | 95.8 | official | |
| 8 | mimo-v2.5-pro | 95.6 | official | |
| 9 | glm-5.2 | 95.4 | official | |
| 9 | gpt-5.4 | 95.4 | official | |
| 11 | grok-4.3 | 94.6 | official | |
| 12 | gemini-3.1-pro-preview | 94.4 | official | |
| 13 | claude-3.7-sonnet | 93.93 | official | |
| 14 | qwen3-235b-a22b | 93.8 | official | |
| 15 | claude-opus-4.5 | 93 | official | |
| 16 | gpt-5 | 92.76 | official | |
| 17 | mimo-v2-pro | 92.2 | official | |
| 18 | gpt-5.2 | 92 | official | |
| 18 | kimi-k2.5 | 92 | official | |
| 20 | hunyuan-large | 91.6 | official | |
| 21 | hunyuan-2.0-thinking | 91.2 | official | |
| 22 | glm-5 | 90.8 | official | |
| 22 | minimax-m3 | 90.8 | official | |
| 24 | grok-4.20 | 90.6 | official | |
| 24 | minimax-m2.7 | 90.6 | official | |
| 26 | claude-sonnet-4 | 90.4 | official | |
| 27 | glm-4.7 | 90.28 | official | |
| 28 | claude-sonnet-4.5 | 90.18 | official | |
| 29 | hy3-preview | 89.6 | official | |
| 30 | gemini-3-pro-preview | 88.4 | official | |
| 31 | doubao-pro-32k | 88.2 | official | |
| 32 | deepseek-v4-pro | 87.8 | official | |
| 33 | minimax-m2 | 86.8 | official | |
| 34 | gpt-5.1 | 86 | official | |
| 35 | doubao-seed-2.0-pro | 85.8 | official | |
| 36 | claude-opus-4.6 | 84 | official | |
| 37 | o4-mini | 83.8 | official | |
| 38 | llama-3.1-nemotron-ultra-253b-v1 | 83.37 | official | |
| 39 | deepseek-v3.2 | 81.2 | official | |
| 40 | glm-4.6 | 80.76 | official | |
| 41 | kimi-k2 | 80.6 | official | |
| 42 | llama-4-maverick | 79.2 | official | |
| 43 | ernie-5.0-thinking | 77.87 | official | |
| 44 | ernie-4.0-turbo | 74.2 | official | |
| 45 | gemini-2.5-pro | 72 | official | |
| 46 | ernie-5.1 | 71.4 | official | |
| 47 | doubao-seed-1.6 | 71.2 | official | |
| 48 | deepseek-v3.1-terminus | 69.4 | official | |
| 49 | doubao-seed-1.8 | 68.84 | official | |
| 50 | ernie-5.0-thinking-preview | 68.8 | official | |
| 51 | ernie-x1.1-preview | 64.24 | official | |
| 52 | doubao-1.5-pro | 64 | official | |
| 53 | deepseek-r1 | 61 | official | |
| 54 | mimo-v2-flash | 60.4 | official | |
| 55 | glm-4.5 | 55.6 | official | |
| 56 | gpt-4o | 48.4 | official | |
| 57 | grok-4 | 41.2 | official | |
| 58 | doubao-1.5-thinking-pro | 40.4 | official | |
| 59 | ernie-x1-turbo | 38.6 | official | |
| 60 | deepseek-v3 | 32 | official | |
| 61 | hunyuan-t1 | 31.4 | official | |
| 62 | minimax-m1-80k | 29.4 | official | |
| 63 | mistral-medium-3 | 25.6 | official |