Evaluation profile
Adversarial Humanities Benchmark (AHB) — Table 5
Weights below are portfolio-specific global index weights.
About this eval
Stylistic robustness of harmful-request refusal across AILuminate-derived hazards under five automated humanities-inspired obfuscation methods.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| child_sexual_exploitationadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:child_sexual_exploitationMeasures unsafe compliance with humanities-obfuscated child sexual exploitation requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.11% | 0.124% |
| defamationadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:defamationMeasures unsafe compliance with humanities-obfuscated defamation requests. | Safety: human_rights_systemic_harm:0.500|truthfulness_honesty:0.500 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.0484% | 0.124% |
| hateadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:hateMeasures unsafe compliance with humanities-obfuscated hateful-content requests. | Safety: fairness_nondiscrimination:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.0905% | 0.124% |
| indiscriminate_weapons_cbrneadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:indiscriminate_weapons_cbrneMeasures unsafe compliance with humanities-obfuscated indiscriminate-weapons and CBRNE requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.0189% | 0.124% |
| intellectual_propertyadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:intellectual_propertyMeasures unsafe compliance with humanities-obfuscated intellectual-property requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.11% | 0.124% |
| non_violent_crimesadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:non_violent_crimesMeasures unsafe compliance with humanities-obfuscated non-violent-crime requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.0189% | 0.124% |
| privacyadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:privacyMeasures unsafe compliance with humanities-obfuscated privacy-violation requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.11% | 0.124% |
| sex_related_crimesadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:sex_related_crimesMeasures unsafe compliance with humanities-obfuscated sex-related-crime requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.11% | 0.124% |
| sexual_contentadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:sexual_contentMeasures unsafe compliance with humanities-obfuscated pornographic-content requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.00755% | 0.124% |
| specialized_adviceadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:specialized_adviceMeasures unsafe compliance with humanities-obfuscated specialized-advice requests while giving each published advice subcategory equal representation. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.0189% | 0.124% |
| suicide_self_harmadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:suicide_self_harmMeasures unsafe compliance with humanities-obfuscated suicide and self-harm requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.11% | 0.124% |
| violent_crimesadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:violent_crimesMeasures unsafe compliance with humanities-obfuscated violent-crime requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmark | Safety: lower · Freedom: higher | 0.11% | 0.124% |
child_sexual_exploitation
Measures unsafe compliance with humanities-obfuscated child sexual exploitation requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.6 | 0 | official | |
| 2 | claude-opus-4.6 | 6.2 | official | |
| 3 | claude-haiku-4.5 | 6.7 | official | |
| 4 | grok-4.1-fast | 14.6 | official | |
| 5 | minimax-m2.7 | 16.7 | official | |
| 6 | qwen3.5-27b | 17 | official | |
| 7 | gpt-5.4-mini | 19.1 | official | |
| 8 | qwen3.5-flash | 21.3 | official | |
| 9 | qwen3.6-plus | 23.9 | official | |
| 10 | gpt-5.4 | 27.1 | official | |
| 11 | qwen3.5-plus | 27.7 | official | |
| 12 | seed-2.0-lite | 31.7 | official | |
| 13 | gpt-oss-20b | 37 | official | |
| 14 | gpt-oss-120b | 38.6 | official | |
| 15 | glm-5.1 | 43.9 | official | |
| 16 | doubao-seed-1.6 | 48.9 | official | |
| 16 | glm-5-turbo | 48.9 | official | |
| 18 | gemini-3.1-pro-preview | 50 | official | |
| 19 | gemma-4-31b-it | 51.6 | official | |
| 20 | llama-4-maverick | 53.2 | official | |
| 21 | llama-4-scout | 55.3 | official | |
| 22 | kimi-k2.5 | 59.5 | official | |
| 23 | deepseek-r1 | 63.8 | official | |
| 23 | qwen3-max | 63.8 | official | |
| 25 | seed-2.0-mini | 64.5 | official | |
| 26 | deepseek-v3.2 | 70 | official | |
| 27 | mistral-small-4 | 70.2 | official | |
| 28 | gemini-3.1-flash-lite | 72.9 | official | |
| 29 | grok-4.20 | 76.6 | official | |
| 30 | mistral-large-3 | 80 | official | |
| 31 | gemini-3-flash-preview | 83.3 | official |
defamation
Measures unsafe compliance with humanities-obfuscated defamation requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 17.4 | official | |
| 2 | claude-sonnet-4.6 | 19.6 | official | |
| 3 | gpt-5.4-mini | 26.1 | official | |
| 4 | minimax-m2.7 | 31.8 | official | |
| 5 | qwen3.5-27b | 33.3 | official | |
| 6 | claude-haiku-4.5 | 37.8 | official | |
| 7 | qwen3.5-flash | 44.4 | official | |
| 8 | gpt-5.4 | 45.7 | official | |
| 9 | qwen3.5-plus | 51.1 | official | |
| 10 | qwen3.6-plus | 55.6 | official | |
| 11 | gpt-oss-120b | 59.1 | official | |
| 12 | llama-4-scout | 64.4 | official | |
| 12 | qwen3-max | 64.4 | official | |
| 14 | mistral-small-4 | 65.2 | official | |
| 15 | gpt-oss-20b | 66.7 | official | |
| 15 | llama-4-maverick | 66.7 | official | |
| 17 | seed-2.0-lite | 69.4 | official | |
| 18 | seed-2.0-mini | 71 | official | |
| 19 | gemma-4-31b-it | 72.7 | official | |
| 20 | doubao-seed-1.6 | 75.6 | official | |
| 21 | deepseek-v3.2 | 77.3 | official | |
| 22 | grok-4.1-fast | 78.3 | official | |
| 23 | mistral-large-3 | 79.1 | official | |
| 24 | kimi-k2.5 | 81.6 | official | |
| 25 | deepseek-r1 | 81.8 | official | |
| 26 | gemini-3.1-flash-lite | 82.6 | official | |
| 27 | grok-4.20 | 84.4 | official | |
| 28 | gemini-3-flash-preview | 84.8 | official | |
| 28 | gemini-3.1-pro-preview | 84.8 | official | |
| 30 | glm-5-turbo | 88.6 | official | |
| 31 | glm-5.1 | 89.2 | official |
hate
Measures unsafe compliance with humanities-obfuscated hateful-content requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.6 | 15.9 | official | |
| 2 | minimax-m2.7 | 23.7 | official | |
| 3 | claude-opus-4.6 | 24.4 | official | |
| 4 | gpt-5.4-mini | 28.9 | official | |
| 5 | qwen3.5-27b | 30.4 | official | |
| 5 | qwen3.5-flash | 30.4 | official | |
| 7 | claude-haiku-4.5 | 31.8 | official | |
| 8 | gpt-5.4 | 35.6 | official | |
| 9 | qwen3.5-plus | 39.1 | official | |
| 10 | gpt-oss-120b | 45.2 | official | |
| 11 | gpt-oss-20b | 47.7 | official | |
| 12 | doubao-seed-1.6 | 47.8 | official | |
| 12 | llama-4-scout | 47.8 | official | |
| 12 | qwen3-max | 47.8 | official | |
| 12 | qwen3.6-plus | 47.8 | official | |
| 16 | llama-4-maverick | 52.2 | official | |
| 17 | seed-2.0-lite | 52.8 | official | |
| 18 | mistral-small-4 | 53.3 | official | |
| 19 | gemma-4-31b-it | 61.1 | official | |
| 20 | deepseek-v3.2 | 64.3 | official | |
| 21 | deepseek-r1 | 64.4 | official | |
| 22 | kimi-k2.5 | 65 | official | |
| 23 | grok-4.1-fast | 65.2 | official | |
| 24 | gemini-3-flash-preview | 66.7 | official | |
| 25 | gemini-3.1-pro-preview | 68.2 | official | |
| 26 | seed-2.0-mini | 70.4 | official | |
| 27 | glm-5-turbo | 71.7 | official | |
| 28 | glm-5.1 | 72.5 | official | |
| 29 | mistral-large-3 | 73.3 | official | |
| 30 | grok-4.20 | 73.9 | official | |
| 31 | gemini-3.1-flash-lite | 76.1 | official |
indiscriminate_weapons_cbrne
Measures unsafe compliance with humanities-obfuscated indiscriminate-weapons and CBRNE requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 0 | official | |
| 1 | claude-sonnet-4.6 | 0 | official | |
| 3 | claude-haiku-4.5 | 8.9 | official | |
| 4 | minimax-m2.7 | 18.6 | official | |
| 5 | gpt-5.4 | 24.4 | official | |
| 6 | qwen3.5-27b | 26.7 | official | |
| 6 | qwen3.6-plus | 26.7 | official | |
| 8 | qwen3.5-plus | 28.9 | official | |
| 9 | qwen3.5-flash | 31.1 | official | |
| 10 | gpt-5.4-mini | 33.3 | official | |
| 10 | grok-4.1-fast | 33.3 | official | |
| 12 | gpt-oss-120b | 60 | official | |
| 13 | gemini-3.1-pro-preview | 61.4 | official | |
| 14 | seed-2.0-mini | 68.8 | official | |
| 15 | seed-2.0-lite | 71.1 | official | |
| 16 | glm-5.1 | 72.2 | official | |
| 17 | gemma-4-31b-it | 73.7 | official | |
| 18 | doubao-seed-1.6 | 73.9 | official | |
| 19 | gpt-oss-20b | 76.7 | official | |
| 20 | grok-4.20 | 80 | official | |
| 20 | llama-4-maverick | 80 | official | |
| 20 | llama-4-scout | 80 | official | |
| 23 | qwen3-max | 80.4 | official | |
| 24 | mistral-small-4 | 82.2 | official | |
| 25 | deepseek-r1 | 84.1 | official | |
| 26 | glm-5-turbo | 86 | official | |
| 27 | gemini-3.1-flash-lite | 86.7 | official | |
| 28 | kimi-k2.5 | 87.9 | official | |
| 29 | gemini-3-flash-preview | 88.9 | official | |
| 30 | mistral-large-3 | 90.5 | official | |
| 31 | deepseek-v3.2 | 90.7 | official |
intellectual_property
Measures unsafe compliance with humanities-obfuscated intellectual-property requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4-mini | 14.9 | official | |
| 2 | claude-sonnet-4.6 | 16.7 | official | |
| 3 | qwen3.5-27b | 25 | official | |
| 4 | gpt-5.4 | 25.5 | official | |
| 5 | claude-opus-4.6 | 29.2 | official | |
| 6 | minimax-m2.7 | 29.3 | official | |
| 7 | claude-haiku-4.5 | 31.2 | official | |
| 8 | qwen3.5-plus | 33.3 | official | |
| 9 | qwen3.5-flash | 37.5 | official | |
| 10 | grok-4.1-fast | 41.3 | official | |
| 11 | qwen3.6-plus | 41.7 | official | |
| 12 | gpt-oss-20b | 46.7 | official | |
| 13 | mistral-small-4 | 51.1 | official | |
| 14 | llama-4-scout | 55.3 | official | |
| 15 | qwen3-max | 56.2 | official | |
| 16 | gpt-oss-120b | 58.7 | official | |
| 17 | llama-4-maverick | 59.6 | official | |
| 18 | grok-4.20 | 66 | official | |
| 19 | gemma-4-31b-it | 67.7 | official | |
| 19 | seed-2.0-mini | 67.7 | official | |
| 21 | gemini-3.1-flash-lite | 68.1 | official | |
| 22 | seed-2.0-lite | 69.4 | official | |
| 23 | glm-5.1 | 69.6 | official | |
| 24 | deepseek-v3.2 | 69.8 | official | |
| 25 | doubao-seed-1.6 | 70.2 | official | |
| 26 | deepseek-r1 | 71.7 | official | |
| 27 | mistral-large-3 | 72.7 | official | |
| 28 | gemini-3.1-pro-preview | 76.6 | official | |
| 29 | glm-5-turbo | 80 | official | |
| 30 | kimi-k2.5 | 80.5 | official | |
| 31 | gemini-3-flash-preview | 81.2 | official |
non_violent_crimes
Measures unsafe compliance with humanities-obfuscated non-violent-crime requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.6 | 8.7 | official | |
| 2 | claude-opus-4.6 | 17.8 | official | |
| 3 | gpt-5.4-mini | 23.9 | official | |
| 4 | claude-haiku-4.5 | 28.3 | official | |
| 5 | qwen3.5-27b | 29.5 | official | |
| 6 | gpt-5.4 | 32.6 | official | |
| 7 | minimax-m2.7 | 36.6 | official | |
| 8 | qwen3.6-plus | 40 | official | |
| 9 | qwen3.5-flash | 43.2 | official | |
| 10 | qwen3.5-plus | 47.7 | official | |
| 11 | grok-4.1-fast | 52.2 | official | |
| 12 | llama-4-scout | 75.6 | official | |
| 13 | seed-2.0-lite | 76.9 | official | |
| 14 | llama-4-maverick | 77.8 | official | |
| 15 | gpt-oss-20b | 78.3 | official | |
| 16 | doubao-seed-1.6 | 81.8 | official | |
| 17 | gemma-4-31b-it | 82.1 | official | |
| 18 | mistral-small-4 | 82.2 | official | |
| 19 | seed-2.0-mini | 83.9 | official | |
| 20 | gpt-oss-120b | 84.1 | official | |
| 21 | kimi-k2.5 | 84.8 | official | |
| 22 | glm-5.1 | 86.5 | official | |
| 23 | gemini-3.1-pro-preview | 86.7 | official | |
| 24 | deepseek-v3.2 | 88.1 | official | |
| 25 | qwen3-max | 88.9 | official | |
| 26 | grok-4.20 | 91.1 | official | |
| 27 | gemini-3.1-flash-lite | 91.3 | official | |
| 28 | deepseek-r1 | 92.9 | official | |
| 29 | gemini-3-flash-preview | 93.5 | official | |
| 29 | glm-5-turbo | 93.5 | official | |
| 31 | mistral-large-3 | 97.7 | official |
privacy
Measures unsafe compliance with humanities-obfuscated privacy-violation requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.6 | 10.9 | official | |
| 2 | claude-opus-4.6 | 23.4 | official | |
| 2 | gpt-5.4-mini | 23.4 | official | |
| 4 | minimax-m2.7 | 31.1 | official | |
| 5 | claude-haiku-4.5 | 31.9 | official | |
| 6 | gpt-5.4 | 33.3 | official | |
| 7 | qwen3.5-27b | 34 | official | |
| 8 | qwen3.5-plus | 42.6 | official | |
| 9 | qwen3.6-plus | 43.8 | official | |
| 10 | qwen3.5-flash | 48.9 | official | |
| 11 | grok-4.1-fast | 53.2 | official | |
| 12 | gpt-oss-120b | 56.5 | official | |
| 13 | seed-2.0-lite | 65 | official | |
| 14 | qwen3-max | 66 | official | |
| 15 | llama-4-maverick | 66.7 | official | |
| 16 | gpt-oss-20b | 68.8 | official | |
| 17 | llama-4-scout | 70.2 | official | |
| 18 | seed-2.0-mini | 70.4 | official | |
| 19 | gemma-4-31b-it | 72.2 | official | |
| 20 | mistral-large-3 | 76.1 | official | |
| 21 | mistral-small-4 | 76.6 | official | |
| 22 | deepseek-r1 | 76.7 | official | |
| 23 | doubao-seed-1.6 | 77.1 | official | |
| 24 | gemini-3.1-flash-lite | 78.7 | official | |
| 24 | gemini-3.1-pro-preview | 78.7 | official | |
| 26 | deepseek-v3.2 | 80 | official | |
| 26 | glm-5-turbo | 80 | official | |
| 26 | kimi-k2.5 | 80 | official | |
| 29 | glm-5.1 | 82.1 | official | |
| 30 | gemini-3-flash-preview | 85.1 | official | |
| 31 | grok-4.20 | 89.6 | official |
sexual_content
Measures unsafe compliance with humanities-obfuscated pornographic-content requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.6 | 6.4 | official | |
| 2 | claude-haiku-4.5 | 17.4 | official | |
| 3 | claude-opus-4.6 | 19.1 | official | |
| 4 | qwen3.5-27b | 22.7 | official | |
| 4 | qwen3.5-flash | 22.7 | official | |
| 6 | gpt-5.4-mini | 27.1 | official | |
| 7 | gpt-5.4 | 29.2 | official | |
| 8 | minimax-m2.7 | 29.3 | official | |
| 9 | qwen3.6-plus | 31.8 | official | |
| 10 | qwen3.5-plus | 34.1 | official | |
| 11 | doubao-seed-1.6 | 35.4 | official | |
| 12 | gpt-oss-20b | 37 | official | |
| 13 | llama-4-scout | 37.5 | official | |
| 14 | seed-2.0-lite | 43.6 | official | |
| 15 | qwen3-max | 45.7 | official | |
| 16 | gpt-oss-120b | 47.7 | official | |
| 17 | mistral-small-4 | 50 | official | |
| 18 | llama-4-maverick | 51.1 | official | |
| 19 | glm-5.1 | 54.8 | official | |
| 20 | gemma-4-31b-it | 55.2 | official | |
| 21 | glm-5-turbo | 55.6 | official | |
| 22 | grok-4.20 | 57.4 | official | |
| 23 | deepseek-v3.2 | 60 | official | |
| 24 | kimi-k2.5 | 60.5 | official | |
| 25 | deepseek-r1 | 61.7 | official | |
| 26 | gemini-3.1-flash-lite | 64.6 | official | |
| 26 | grok-4.1-fast | 64.6 | official | |
| 28 | gemini-3-flash-preview | 66.7 | official | |
| 29 | mistral-large-3 | 67.4 | official | |
| 30 | gemini-3.1-pro-preview | 68.8 | official | |
| 31 | seed-2.0-mini | 72.7 | official |
specialized_advice
Measures unsafe compliance with humanities-obfuscated specialized-advice requests while giving each published advice subcategory equal representation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.6 | 12.6 | official | |
| 2 | claude-opus-4.6 | 20.62 | official | |
| 3 | gpt-5.4-mini | 22.48 | official | |
| 4 | gpt-5.4 | 28.32 | official | |
| 5 | claude-haiku-4.5 | 32.67 | official | |
| 6 | minimax-m2.7 | 33.92 | official | |
| 7 | qwen3.5-27b | 35.95 | official | |
| 8 | qwen3.5-flash | 42.38 | official | |
| 9 | qwen3.5-plus | 42.48 | official | |
| 10 | qwen3.6-plus | 46.73 | official | |
| 11 | gpt-oss-120b | 55.73 | official | |
| 12 | llama-4-scout | 56.55 | official | |
| 13 | llama-4-maverick | 60.1 | official | |
| 14 | grok-4.1-fast | 60.92 | official | |
| 15 | qwen3-max | 61.38 | official | |
| 16 | mistral-small-4 | 61.83 | official | |
| 17 | gpt-oss-20b | 62.6 | official | |
| 18 | deepseek-r1 | 66.05 | official | |
| 19 | gemma-4-31b-it | 68 | official | |
| 20 | kimi-k2.5 | 69.05 | official | |
| 21 | deepseek-v3.2 | 69.38 | official | |
| 22 | seed-2.0-lite | 70.45 | official | |
| 23 | seed-2.0-mini | 70.67 | official | |
| 24 | gemini-3.1-flash-lite | 71.47 | official | |
| 25 | doubao-seed-1.6 | 72.38 | official | |
| 26 | glm-5-turbo | 73.2 | official | |
| 27 | glm-5.1 | 73.85 | official | |
| 28 | gemini-3.1-pro-preview | 74.5 | official | |
| 29 | grok-4.20 | 75.33 | official | |
| 30 | gemini-3-flash-preview | 76.53 | official | |
| 31 | mistral-large-3 | 77.1 | official |
suicide_self_harm
Measures unsafe compliance with humanities-obfuscated suicide and self-harm requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 4.2 | official | |
| 1 | claude-sonnet-4.6 | 4.2 | official | |
| 3 | gpt-5.4-mini | 8.5 | official | |
| 4 | claude-haiku-4.5 | 18.8 | official | |
| 5 | qwen3.5-27b | 19.1 | official | |
| 6 | gpt-5.4 | 22.9 | official | |
| 7 | minimax-m2.7 | 26.8 | official | |
| 8 | qwen3.5-flash | 27.7 | official | |
| 9 | qwen3.5-plus | 29.8 | official | |
| 10 | qwen3.6-plus | 34 | official | |
| 11 | gpt-oss-120b | 44.2 | official | |
| 12 | gpt-oss-20b | 45.8 | official | |
| 13 | grok-4.1-fast | 55.3 | official | |
| 14 | glm-5.1 | 58.3 | official | |
| 14 | llama-4-maverick | 58.3 | official | |
| 16 | qwen3-max | 60.4 | official | |
| 17 | gemini-3.1-pro-preview | 60.9 | official | |
| 18 | gemma-4-31b-it | 61.8 | official | |
| 18 | seed-2.0-lite | 61.8 | official | |
| 20 | llama-4-scout | 62.5 | official | |
| 20 | mistral-small-4 | 62.5 | official | |
| 22 | deepseek-v3.2 | 65.1 | official | |
| 23 | doubao-seed-1.6 | 66.7 | official | |
| 23 | kimi-k2.5 | 66.7 | official | |
| 25 | deepseek-r1 | 67.4 | official | |
| 26 | glm-5-turbo | 69.6 | official | |
| 27 | grok-4.20 | 72.9 | official | |
| 28 | gemini-3-flash-preview | 77.1 | official | |
| 29 | seed-2.0-mini | 78.8 | official | |
| 30 | gemini-3.1-flash-lite | 80.9 | official | |
| 31 | mistral-large-3 | 84.4 | official |
violent_crimes
Measures unsafe compliance with humanities-obfuscated violent-crime requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 2.2 | official | |
| 1 | claude-sonnet-4.6 | 2.2 | official | |
| 3 | claude-haiku-4.5 | 23.9 | official | |
| 3 | gpt-5.4-mini | 23.9 | official | |
| 5 | gpt-5.4 | 26.7 | official | |
| 6 | minimax-m2.7 | 35 | official | |
| 7 | qwen3.5-27b | 37 | official | |
| 7 | qwen3.5-flash | 37 | official | |
| 9 | qwen3.5-plus | 41.3 | official | |
| 10 | qwen3.6-plus | 52.2 | official | |
| 11 | seed-2.0-lite | 55.6 | official | |
| 12 | gpt-oss-120b | 56.8 | official | |
| 13 | grok-4.1-fast | 60 | official | |
| 14 | doubao-seed-1.6 | 64.4 | official | |
| 15 | mistral-small-4 | 65.2 | official | |
| 16 | gpt-oss-20b | 69 | official | |
| 17 | seed-2.0-mini | 69.7 | official | |
| 18 | gemma-4-31b-it | 75.8 | official | |
| 19 | llama-4-maverick | 76.1 | official | |
| 19 | qwen3-max | 76.1 | official | |
| 21 | grok-4.20 | 78.3 | official | |
| 22 | llama-4-scout | 80.4 | official | |
| 23 | deepseek-r1 | 81.8 | official | |
| 24 | glm-5.1 | 82.9 | official | |
| 25 | mistral-large-3 | 83.3 | official | |
| 26 | gemini-3.1-pro-preview | 84.1 | official | |
| 27 | gemini-3-flash-preview | 88.9 | official | |
| 27 | gemini-3.1-flash-lite | 88.9 | official | |
| 29 | deepseek-v3.2 | 90.2 | official | |
| 30 | kimi-k2.5 | 90.3 | official | |
| 31 | glm-5-turbo | 93 | official |