← Evals

Evaluation profile

Adversarial Humanities Benchmark (AHB) — Table 5

12sub-evals
0.861%Safety weight
1.49%Freedom weight
4components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Stylistic robustness of harmful-request refusal across AILuminate-derived hazards under five automated humanities-inspired obfuscation methods.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
child_sexual_exploitationadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:child_sexual_exploitationMeasures unsafe compliance with humanities-obfuscated child sexual exploitation requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.11%0.124%
defamationadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:defamationMeasures unsafe compliance with humanities-obfuscated defamation requests.Safety: human_rights_systemic_harm:0.500|truthfulness_honesty:0.500 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.0484%0.124%
hateadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:hateMeasures unsafe compliance with humanities-obfuscated hateful-content requests.Safety: fairness_nondiscrimination:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.0905%0.124%
indiscriminate_weapons_cbrneadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:indiscriminate_weapons_cbrneMeasures unsafe compliance with humanities-obfuscated indiscriminate-weapons and CBRNE requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.0189%0.124%
intellectual_propertyadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:intellectual_propertyMeasures unsafe compliance with humanities-obfuscated intellectual-property requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.11%0.124%
non_violent_crimesadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:non_violent_crimesMeasures unsafe compliance with humanities-obfuscated non-violent-crime requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.0189%0.124%
privacyadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:privacyMeasures unsafe compliance with humanities-obfuscated privacy-violation requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.11%0.124%
sex_related_crimesadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:sex_related_crimesMeasures unsafe compliance with humanities-obfuscated sex-related-crime requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.11%0.124%
sexual_contentadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:sexual_contentMeasures unsafe compliance with humanities-obfuscated pornographic-content requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.00755%0.124%
specialized_adviceadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:specialized_adviceMeasures unsafe compliance with humanities-obfuscated specialized-advice requests while giving each published advice subcategory equal representation.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.0189%0.124%
suicide_self_harmadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:suicide_self_harmMeasures unsafe compliance with humanities-obfuscated suicide and self-harm requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.11%0.124%
violent_crimesadversarial-humanities-benchmark/adversarial-humanities-benchmark.csv:violent_crimesMeasures unsafe compliance with humanities-obfuscated violent-crime requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000adversarial-humanities-benchmarkSafety: lower · Freedom: higher0.11%0.124%

child_sexual_exploitation

Measures unsafe compliance with humanities-obfuscated child sexual exploitation requests.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.60official
2claude-opus-4.66.2official
3claude-haiku-4.56.7official
4grok-4.1-fast14.6official
5minimax-m2.716.7official
6qwen3.5-27b17official
7gpt-5.4-mini19.1official
8qwen3.5-flash21.3official
9qwen3.6-plus23.9official
10gpt-5.427.1official
11qwen3.5-plus27.7official
12seed-2.0-lite31.7official
13gpt-oss-20b37official
14gpt-oss-120b38.6official
15glm-5.143.9official
16doubao-seed-1.648.9official
16glm-5-turbo48.9official
18gemini-3.1-pro-preview50official
19gemma-4-31b-it51.6official
20llama-4-maverick53.2official
21llama-4-scout55.3official
22kimi-k2.559.5official
23deepseek-r163.8official
23qwen3-max63.8official
25seed-2.0-mini64.5official
26deepseek-v3.270official
27mistral-small-470.2official
28gemini-3.1-flash-lite72.9official
29grok-4.2076.6official
30mistral-large-380official
31gemini-3-flash-preview83.3official

defamation

Measures unsafe compliance with humanities-obfuscated defamation requests.

RankModelValueRelative performanceProvenance
1claude-opus-4.617.4official
2claude-sonnet-4.619.6official
3gpt-5.4-mini26.1official
4minimax-m2.731.8official
5qwen3.5-27b33.3official
6claude-haiku-4.537.8official
7qwen3.5-flash44.4official
8gpt-5.445.7official
9qwen3.5-plus51.1official
10qwen3.6-plus55.6official
11gpt-oss-120b59.1official
12llama-4-scout64.4official
12qwen3-max64.4official
14mistral-small-465.2official
15gpt-oss-20b66.7official
15llama-4-maverick66.7official
17seed-2.0-lite69.4official
18seed-2.0-mini71official
19gemma-4-31b-it72.7official
20doubao-seed-1.675.6official
21deepseek-v3.277.3official
22grok-4.1-fast78.3official
23mistral-large-379.1official
24kimi-k2.581.6official
25deepseek-r181.8official
26gemini-3.1-flash-lite82.6official
27grok-4.2084.4official
28gemini-3-flash-preview84.8official
28gemini-3.1-pro-preview84.8official
30glm-5-turbo88.6official
31glm-5.189.2official

hate

Measures unsafe compliance with humanities-obfuscated hateful-content requests.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.615.9official
2minimax-m2.723.7official
3claude-opus-4.624.4official
4gpt-5.4-mini28.9official
5qwen3.5-27b30.4official
5qwen3.5-flash30.4official
7claude-haiku-4.531.8official
8gpt-5.435.6official
9qwen3.5-plus39.1official
10gpt-oss-120b45.2official
11gpt-oss-20b47.7official
12doubao-seed-1.647.8official
12llama-4-scout47.8official
12qwen3-max47.8official
12qwen3.6-plus47.8official
16llama-4-maverick52.2official
17seed-2.0-lite52.8official
18mistral-small-453.3official
19gemma-4-31b-it61.1official
20deepseek-v3.264.3official
21deepseek-r164.4official
22kimi-k2.565official
23grok-4.1-fast65.2official
24gemini-3-flash-preview66.7official
25gemini-3.1-pro-preview68.2official
26seed-2.0-mini70.4official
27glm-5-turbo71.7official
28glm-5.172.5official
29mistral-large-373.3official
30grok-4.2073.9official
31gemini-3.1-flash-lite76.1official

indiscriminate_weapons_cbrne

Measures unsafe compliance with humanities-obfuscated indiscriminate-weapons and CBRNE requests.

RankModelValueRelative performanceProvenance
1claude-opus-4.60official
1claude-sonnet-4.60official
3claude-haiku-4.58.9official
4minimax-m2.718.6official
5gpt-5.424.4official
6qwen3.5-27b26.7official
6qwen3.6-plus26.7official
8qwen3.5-plus28.9official
9qwen3.5-flash31.1official
10gpt-5.4-mini33.3official
10grok-4.1-fast33.3official
12gpt-oss-120b60official
13gemini-3.1-pro-preview61.4official
14seed-2.0-mini68.8official
15seed-2.0-lite71.1official
16glm-5.172.2official
17gemma-4-31b-it73.7official
18doubao-seed-1.673.9official
19gpt-oss-20b76.7official
20grok-4.2080official
20llama-4-maverick80official
20llama-4-scout80official
23qwen3-max80.4official
24mistral-small-482.2official
25deepseek-r184.1official
26glm-5-turbo86official
27gemini-3.1-flash-lite86.7official
28kimi-k2.587.9official
29gemini-3-flash-preview88.9official
30mistral-large-390.5official
31deepseek-v3.290.7official

intellectual_property

Measures unsafe compliance with humanities-obfuscated intellectual-property requests.

RankModelValueRelative performanceProvenance
1gpt-5.4-mini14.9official
2claude-sonnet-4.616.7official
3qwen3.5-27b25official
4gpt-5.425.5official
5claude-opus-4.629.2official
6minimax-m2.729.3official
7claude-haiku-4.531.2official
8qwen3.5-plus33.3official
9qwen3.5-flash37.5official
10grok-4.1-fast41.3official
11qwen3.6-plus41.7official
12gpt-oss-20b46.7official
13mistral-small-451.1official
14llama-4-scout55.3official
15qwen3-max56.2official
16gpt-oss-120b58.7official
17llama-4-maverick59.6official
18grok-4.2066official
19gemma-4-31b-it67.7official
19seed-2.0-mini67.7official
21gemini-3.1-flash-lite68.1official
22seed-2.0-lite69.4official
23glm-5.169.6official
24deepseek-v3.269.8official
25doubao-seed-1.670.2official
26deepseek-r171.7official
27mistral-large-372.7official
28gemini-3.1-pro-preview76.6official
29glm-5-turbo80official
30kimi-k2.580.5official
31gemini-3-flash-preview81.2official

non_violent_crimes

Measures unsafe compliance with humanities-obfuscated non-violent-crime requests.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.68.7official
2claude-opus-4.617.8official
3gpt-5.4-mini23.9official
4claude-haiku-4.528.3official
5qwen3.5-27b29.5official
6gpt-5.432.6official
7minimax-m2.736.6official
8qwen3.6-plus40official
9qwen3.5-flash43.2official
10qwen3.5-plus47.7official
11grok-4.1-fast52.2official
12llama-4-scout75.6official
13seed-2.0-lite76.9official
14llama-4-maverick77.8official
15gpt-oss-20b78.3official
16doubao-seed-1.681.8official
17gemma-4-31b-it82.1official
18mistral-small-482.2official
19seed-2.0-mini83.9official
20gpt-oss-120b84.1official
21kimi-k2.584.8official
22glm-5.186.5official
23gemini-3.1-pro-preview86.7official
24deepseek-v3.288.1official
25qwen3-max88.9official
26grok-4.2091.1official
27gemini-3.1-flash-lite91.3official
28deepseek-r192.9official
29gemini-3-flash-preview93.5official
29glm-5-turbo93.5official
31mistral-large-397.7official

privacy

Measures unsafe compliance with humanities-obfuscated privacy-violation requests.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.610.9official
2claude-opus-4.623.4official
2gpt-5.4-mini23.4official
4minimax-m2.731.1official
5claude-haiku-4.531.9official
6gpt-5.433.3official
7qwen3.5-27b34official
8qwen3.5-plus42.6official
9qwen3.6-plus43.8official
10qwen3.5-flash48.9official
11grok-4.1-fast53.2official
12gpt-oss-120b56.5official
13seed-2.0-lite65official
14qwen3-max66official
15llama-4-maverick66.7official
16gpt-oss-20b68.8official
17llama-4-scout70.2official
18seed-2.0-mini70.4official
19gemma-4-31b-it72.2official
20mistral-large-376.1official
21mistral-small-476.6official
22deepseek-r176.7official
23doubao-seed-1.677.1official
24gemini-3.1-flash-lite78.7official
24gemini-3.1-pro-preview78.7official
26deepseek-v3.280official
26glm-5-turbo80official
26kimi-k2.580official
29glm-5.182.1official
30gemini-3-flash-preview85.1official
31grok-4.2089.6official

sexual_content

Measures unsafe compliance with humanities-obfuscated pornographic-content requests.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.66.4official
2claude-haiku-4.517.4official
3claude-opus-4.619.1official
4qwen3.5-27b22.7official
4qwen3.5-flash22.7official
6gpt-5.4-mini27.1official
7gpt-5.429.2official
8minimax-m2.729.3official
9qwen3.6-plus31.8official
10qwen3.5-plus34.1official
11doubao-seed-1.635.4official
12gpt-oss-20b37official
13llama-4-scout37.5official
14seed-2.0-lite43.6official
15qwen3-max45.7official
16gpt-oss-120b47.7official
17mistral-small-450official
18llama-4-maverick51.1official
19glm-5.154.8official
20gemma-4-31b-it55.2official
21glm-5-turbo55.6official
22grok-4.2057.4official
23deepseek-v3.260official
24kimi-k2.560.5official
25deepseek-r161.7official
26gemini-3.1-flash-lite64.6official
26grok-4.1-fast64.6official
28gemini-3-flash-preview66.7official
29mistral-large-367.4official
30gemini-3.1-pro-preview68.8official
31seed-2.0-mini72.7official

specialized_advice

Measures unsafe compliance with humanities-obfuscated specialized-advice requests while giving each published advice subcategory equal representation.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.612.6official
2claude-opus-4.620.62official
3gpt-5.4-mini22.48official
4gpt-5.428.32official
5claude-haiku-4.532.67official
6minimax-m2.733.92official
7qwen3.5-27b35.95official
8qwen3.5-flash42.38official
9qwen3.5-plus42.48official
10qwen3.6-plus46.73official
11gpt-oss-120b55.73official
12llama-4-scout56.55official
13llama-4-maverick60.1official
14grok-4.1-fast60.92official
15qwen3-max61.38official
16mistral-small-461.83official
17gpt-oss-20b62.6official
18deepseek-r166.05official
19gemma-4-31b-it68official
20kimi-k2.569.05official
21deepseek-v3.269.38official
22seed-2.0-lite70.45official
23seed-2.0-mini70.67official
24gemini-3.1-flash-lite71.47official
25doubao-seed-1.672.38official
26glm-5-turbo73.2official
27glm-5.173.85official
28gemini-3.1-pro-preview74.5official
29grok-4.2075.33official
30gemini-3-flash-preview76.53official
31mistral-large-377.1official

suicide_self_harm

Measures unsafe compliance with humanities-obfuscated suicide and self-harm requests.

RankModelValueRelative performanceProvenance
1claude-opus-4.64.2official
1claude-sonnet-4.64.2official
3gpt-5.4-mini8.5official
4claude-haiku-4.518.8official
5qwen3.5-27b19.1official
6gpt-5.422.9official
7minimax-m2.726.8official
8qwen3.5-flash27.7official
9qwen3.5-plus29.8official
10qwen3.6-plus34official
11gpt-oss-120b44.2official
12gpt-oss-20b45.8official
13grok-4.1-fast55.3official
14glm-5.158.3official
14llama-4-maverick58.3official
16qwen3-max60.4official
17gemini-3.1-pro-preview60.9official
18gemma-4-31b-it61.8official
18seed-2.0-lite61.8official
20llama-4-scout62.5official
20mistral-small-462.5official
22deepseek-v3.265.1official
23doubao-seed-1.666.7official
23kimi-k2.566.7official
25deepseek-r167.4official
26glm-5-turbo69.6official
27grok-4.2072.9official
28gemini-3-flash-preview77.1official
29seed-2.0-mini78.8official
30gemini-3.1-flash-lite80.9official
31mistral-large-384.4official

violent_crimes

Measures unsafe compliance with humanities-obfuscated violent-crime requests.

RankModelValueRelative performanceProvenance
1claude-opus-4.62.2official
1claude-sonnet-4.62.2official
3claude-haiku-4.523.9official
3gpt-5.4-mini23.9official
5gpt-5.426.7official
6minimax-m2.735official
7qwen3.5-27b37official
7qwen3.5-flash37official
9qwen3.5-plus41.3official
10qwen3.6-plus52.2official
11seed-2.0-lite55.6official
12gpt-oss-120b56.8official
13grok-4.1-fast60official
14doubao-seed-1.664.4official
15mistral-small-465.2official
16gpt-oss-20b69official
17seed-2.0-mini69.7official
18gemma-4-31b-it75.8official
19llama-4-maverick76.1official
19qwen3-max76.1official
21grok-4.2078.3official
22llama-4-scout80.4official
23deepseek-r181.8official
24glm-5.182.9official
25mistral-large-383.3official
26gemini-3.1-pro-preview84.1official
27gemini-3-flash-preview88.9official
27gemini-3.1-flash-lite88.9official
29deepseek-v3.290.2official
30kimi-k2.590.3official
31glm-5-turbo93official