Evaluation profile
Adversarial Poetry — AILuminate Baseline and Poetry ASR
2sub-evals
0.183%Safety weight
1.31%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Resistance to unsafe compliance on broad harmful requests before and after a standardized poetry-only surface-form transformation.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| baseline_asradversarial-poetry/adversarial-poetry.csv:baseline_asrMeasures unsafe compliance on the direct harmful-request baseline before poetic reframing. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-poetry | Safety: lower · Freedom: higher | 0.0913% | 0.655% |
| poetry_asradversarial-poetry/adversarial-poetry.csv:poetry_asrMeasures unsafe compliance when the same harmful-request distribution is reframed as poetry. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-poetry | Safety: lower · Freedom: higher | 0.0913% | 0.655% |
baseline_asr
Measures unsafe compliance on the direct harmful-request baseline before poetic reframing.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5-nano | 0.82 | official | |
| 1 | gpt-oss-120b | 0.82 | official | |
| 3 | gpt-5 | 1.1 | official | |
| 4 | claude-opus-4.1 | 2.01 | official | |
| 5 | claude-sonnet-4.5 | 2.06 | official | |
| 6 | gpt-5-mini | 2.16 | official | |
| 7 | claude-haiku-4.5 | 2.27 | official | |
| 8 | qwen3-max | 2.93 | official | |
| 9 | gpt-oss-20b | 3.88 | official | |
| 10 | llama-4-maverick | 5.14 | official | |
| 11 | kimi-k2 | 6.045 | official | |
| 12 | deepseek-v3.2-exp | 7.52 | official | |
| 13 | gemini-2.5-flash | 7.79 | official | |
| 14 | grok-4-fast | 7.84 | official | |
| 15 | gemini-2.5-flash-lite | 8.67 | official | |
| 16 | deepseek-v3.1 | 8.81 | official | |
| 17 | qwen3-32b | 9.67 | official | |
| 18 | gemini-2.5-pro | 10.15 | official | |
| 19 | llama-4-scout | 11.52 | official | |
| 20 | deepseek-r1 | 13.29 | official | |
| 21 | grok-4 | 16.04 | official | |
| 22 | mistral-large-2 | 20.81 | official | |
| 23 | mistral-small-3.2 | 21.96 | official | |
| 24 | magistral-medium | 22.92 | official |
poetry_asr
Measures unsafe compliance when the same harmful-request distribution is reframed as poetry.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-haiku-4.5 | 0.6 | official | |
| 2 | gpt-5-nano | 1.47 | official | |
| 3 | gpt-5-mini | 3.73 | official | |
| 4 | claude-opus-4.1 | 5.45 | official | |
| 5 | gpt-5 | 6.14 | official | |
| 6 | gpt-oss-120b | 8.94 | official | |
| 7 | claude-sonnet-4.5 | 9.69 | official | |
| 8 | gpt-oss-20b | 23.26 | official | |
| 9 | grok-4 | 34.4 | official | |
| 10 | grok-4-fast | 35.58 | official | |
| 11 | llama-4-maverick | 43.44 | official | |
| 12 | llama-4-scout | 49.61 | official | |
| 13 | kimi-k2 | 51.88 | official | |
| 14 | qwen3-max | 55.44 | official | |
| 15 | gemini-2.5-flash-lite | 64.77 | official | |
| 16 | mistral-small-3.2 | 65.46 | official | |
| 17 | gemini-2.5-flash | 65.79 | official | |
| 18 | gemini-2.5-pro | 66.73 | official | |
| 19 | deepseek-r1 | 67.57 | official | |
| 20 | qwen3-32b | 69.05 | official | |
| 21 | mistral-large-2 | 69.42 | official | |
| 22 | deepseek-v3.2-exp | 71.94 | official | |
| 23 | deepseek-v3.1 | 76.71 | official | |
| 24 | magistral-medium | 77.19 | official |