Evaluation profile
Adversarial Poetry Refusal (AHB self-run)
Weights below are portfolio-specific global index weights.
About this eval
Measures attack-success rate on the 600 adversarial-poetry prompts published in the Hugging Face AHB repository; lower values indicate stronger refusal of harmful requests.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| child_sexual_exploitationahb-poetry-self-run/ahb-poetry-self-run.csv:child_sexual_exploitationMeasures unsafe compliance with poetry-obfuscated child-sexual-exploitation requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0557% | 0.0631% |
| defamationahb-poetry-self-run/ahb-poetry-self-run.csv:defamationMeasures unsafe compliance with poetry-obfuscated defamatory requests. | Safety: human_rights_systemic_harm:0.500|truthfulness_honesty:0.500 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0246% | 0.0631% |
| hateahb-poetry-self-run/ahb-poetry-self-run.csv:hateMeasures unsafe compliance with poetry-obfuscated hateful requests. | Safety: fairness_nondiscrimination:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.046% | 0.0631% |
| indiscriminate_weapons_cbrneahb-poetry-self-run/ahb-poetry-self-run.csv:indiscriminate_weapons_cbrneMeasures unsafe compliance with poetry-obfuscated indiscriminate-weapons and CBRNE requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.00959% | 0.0631% |
| intellectual_propertyahb-poetry-self-run/ahb-poetry-self-run.csv:intellectual_propertyMeasures unsafe compliance with poetry-obfuscated intellectual-property requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0557% | 0.0631% |
| non_violent_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:non_violent_crimesMeasures unsafe compliance with poetry-obfuscated non-violent-crime requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.00959% | 0.0631% |
| privacyahb-poetry-self-run/ahb-poetry-self-run.csv:privacyMeasures unsafe compliance with poetry-obfuscated privacy-violation requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0557% | 0.0631% |
| sex_related_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:sex_related_crimesMeasures unsafe compliance with poetry-obfuscated sex-related-crime requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0557% | 0.0631% |
| sexual_contentahb-poetry-self-run/ahb-poetry-self-run.csv:sexual_contentMeasures unsafe compliance with poetry-obfuscated sexual-content requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.00384% | 0.0631% |
| specialized_adviceahb-poetry-self-run/ahb-poetry-self-run.csv:specialized_adviceMeasures unsafe compliance with poetry-obfuscated specialized-advice requests using the known item-level denominators. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.00959% | 0.0631% |
| suicide_self_harmahb-poetry-self-run/ahb-poetry-self-run.csv:suicide_self_harmMeasures unsafe compliance with poetry-obfuscated suicide and self-harm requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0557% | 0.0631% |
| violent_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:violent_crimesMeasures unsafe compliance with poetry-obfuscated violent-crime requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0557% | 0.0631% |
child_sexual_exploitation
Measures unsafe compliance with poetry-obfuscated child-sexual-exploitation requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | glm-5.3-flash | 1.449 | self run | |
| 1 | gpt-5.6-luna | 1.449 | self run | |
| 1 | mercury-2 | 1.449 | self run | |
| 4 | granite-4.2-8b | 2.899 | self run | |
| 5 | longcat-2.0 | 4.348 | self run | |
| 6 | laguna-s-2.1-poolside | 7.971 | self run | |
| 6 | nemotron-3.5-lightning | 7.971 | self run | |
| 8 | solar-pro-4 | 31.88 | self run |
defamation
Measures unsafe compliance with poetry-obfuscated defamatory requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | mercury-2 | 2.174 | self run | |
| 2 | glm-5.3-flash | 3.623 | self run | |
| 2 | granite-4.2-8b | 3.623 | self run | |
| 4 | laguna-s-2.1-poolside | 7.971 | self run | |
| 5 | gpt-5.6-luna | 9.42 | self run | |
| 6 | longcat-2.0 | 17.39 | self run | |
| 7 | nemotron-3.5-lightning | 19.57 | self run | |
| 8 | solar-pro-4 | 55.07 | self run |
hate
Measures unsafe compliance with poetry-obfuscated hateful requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | glm-5.3-flash | 3.623 | self run | |
| 2 | granite-4.2-8b | 5.072 | self run | |
| 3 | gpt-5.6-luna | 5.797 | self run | |
| 3 | longcat-2.0 | 5.797 | self run | |
| 3 | mercury-2 | 5.797 | self run | |
| 6 | laguna-s-2.1-poolside | 7.246 | self run | |
| 7 | nemotron-3.5-lightning | 12.32 | self run | |
| 8 | solar-pro-4 | 43.48 | self run |
indiscriminate_weapons_cbrne
Measures unsafe compliance with poetry-obfuscated indiscriminate-weapons and CBRNE requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0 | self run | |
| 1 | mercury-2 | 0 | self run | |
| 3 | glm-5.3-flash | 1.449 | self run | |
| 4 | longcat-2.0 | 7.971 | self run | |
| 5 | laguna-s-2.1-poolside | 9.42 | self run | |
| 6 | granite-4.2-8b | 10.87 | self run | |
| 7 | nemotron-3.5-lightning | 12.32 | self run | |
| 8 | solar-pro-4 | 34.06 | self run |
intellectual_property
Measures unsafe compliance with poetry-obfuscated intellectual-property requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0 | self run | |
| 2 | mercury-2 | 2.174 | self run | |
| 3 | granite-4.2-8b | 3.623 | self run | |
| 4 | glm-5.3-flash | 8.696 | self run | |
| 5 | laguna-s-2.1-poolside | 10.14 | self run | |
| 6 | longcat-2.0 | 10.87 | self run | |
| 7 | nemotron-3.5-lightning | 21.74 | self run | |
| 8 | solar-pro-4 | 27.54 | self run |
non_violent_crimes
Measures unsafe compliance with poetry-obfuscated non-violent-crime requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0.7246 | self run | |
| 2 | mercury-2 | 2.174 | self run | |
| 3 | glm-5.3-flash | 4.348 | self run | |
| 3 | laguna-s-2.1-poolside | 4.348 | self run | |
| 5 | granite-4.2-8b | 8.696 | self run | |
| 6 | longcat-2.0 | 10.14 | self run | |
| 7 | nemotron-3.5-lightning | 18.12 | self run | |
| 8 | solar-pro-4 | 34.78 | self run |
privacy
Measures unsafe compliance with poetry-obfuscated privacy-violation requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 5.797 | self run | |
| 2 | glm-5.3-flash | 7.246 | self run | |
| 3 | mercury-2 | 9.42 | self run | |
| 4 | granite-4.2-8b | 13.77 | self run | |
| 5 | laguna-s-2.1-poolside | 15.94 | self run | |
| 6 | longcat-2.0 | 18.84 | self run | |
| 7 | nemotron-3.5-lightning | 27.54 | self run | |
| 8 | solar-pro-4 | 43.48 | self run |
sexual_content
Measures unsafe compliance with poetry-obfuscated sexual-content requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | granite-4.2-8b | 4.545 | self run | |
| 1 | mercury-2 | 4.545 | self run | |
| 3 | gpt-5.6-luna | 5.303 | self run | |
| 4 | laguna-s-2.1-poolside | 6.818 | self run | |
| 5 | glm-5.3-flash | 7.576 | self run | |
| 6 | nemotron-3.5-lightning | 12.88 | self run | |
| 7 | longcat-2.0 | 15.15 | self run | |
| 8 | solar-pro-4 | 25 | self run |
specialized_advice
Measures unsafe compliance with poetry-obfuscated specialized-advice requests using the known item-level denominators.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 2.333 | self run | |
| 1 | mercury-2 | 2.333 | self run | |
| 3 | glm-5.3-flash | 3 | self run | |
| 4 | granite-4.2-8b | 4.667 | self run | |
| 4 | laguna-s-2.1-poolside | 4.667 | self run | |
| 6 | longcat-2.0 | 6.333 | self run | |
| 7 | nemotron-3.5-lightning | 12.67 | self run | |
| 8 | solar-pro-4 | 29.67 | self run |
suicide_self_harm
Measures unsafe compliance with poetry-obfuscated suicide and self-harm requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 1.515 | self run | |
| 2 | mercury-2 | 3.788 | self run | |
| 3 | laguna-s-2.1-poolside | 6.061 | self run | |
| 4 | granite-4.2-8b | 6.818 | self run | |
| 5 | glm-5.3-flash | 7.576 | self run | |
| 6 | longcat-2.0 | 9.848 | self run | |
| 7 | nemotron-3.5-lightning | 15.15 | self run | |
| 8 | solar-pro-4 | 37.12 | self run |
violent_crimes
Measures unsafe compliance with poetry-obfuscated violent-crime requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 2.273 | self run | |
| 1 | mercury-2 | 2.273 | self run | |
| 3 | glm-5.3-flash | 5.303 | self run | |
| 4 | laguna-s-2.1-poolside | 7.576 | self run | |
| 5 | granite-4.2-8b | 8.333 | self run | |
| 6 | longcat-2.0 | 14.39 | self run | |
| 7 | nemotron-3.5-lightning | 18.94 | self run | |
| 8 | solar-pro-4 | 47.73 | self run |