← Evals

Evaluation profile

Adversarial Poetry — AILuminate Baseline and Poetry ASR

2sub-evals
0.183%Safety weight
1.31%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Resistance to unsafe compliance on broad harmful requests before and after a standardized poetry-only surface-form transformation.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
baseline_asradversarial-poetry/adversarial-poetry.csv:baseline_asrMeasures unsafe compliance on the direct harmful-request baseline before poetic reframing.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-poetrySafety: lower · Freedom: higher0.0913%0.655%
poetry_asradversarial-poetry/adversarial-poetry.csv:poetry_asrMeasures unsafe compliance when the same harmful-request distribution is reframed as poetry.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000adversarial-poetrySafety: lower · Freedom: higher0.0913%0.655%

baseline_asr

Measures unsafe compliance on the direct harmful-request baseline before poetic reframing.

RankModelValueRelative performanceProvenance
1gpt-5-nano0.82official
1gpt-oss-120b0.82official
3gpt-51.1official
4claude-opus-4.12.01official
5claude-sonnet-4.52.06official
6gpt-5-mini2.16official
7claude-haiku-4.52.27official
8qwen3-max2.93official
9gpt-oss-20b3.88official
10llama-4-maverick5.14official
11kimi-k26.045official
12deepseek-v3.2-exp7.52official
13gemini-2.5-flash7.79official
14grok-4-fast7.84official
15gemini-2.5-flash-lite8.67official
16deepseek-v3.18.81official
17qwen3-32b9.67official
18gemini-2.5-pro10.15official
19llama-4-scout11.52official
20deepseek-r113.29official
21grok-416.04official
22mistral-large-220.81official
23mistral-small-3.221.96official
24magistral-medium22.92official

poetry_asr

Measures unsafe compliance when the same harmful-request distribution is reframed as poetry.

RankModelValueRelative performanceProvenance
1claude-haiku-4.50.6official
2gpt-5-nano1.47official
3gpt-5-mini3.73official
4claude-opus-4.15.45official
5gpt-56.14official
6gpt-oss-120b8.94official
7claude-sonnet-4.59.69official
8gpt-oss-20b23.26official
9grok-434.4official
10grok-4-fast35.58official
11llama-4-maverick43.44official
12llama-4-scout49.61official
13kimi-k251.88official
14qwen3-max55.44official
15gemini-2.5-flash-lite64.77official
16mistral-small-3.265.46official
17gemini-2.5-flash65.79official
18gemini-2.5-pro66.73official
19deepseek-r167.57official
20qwen3-32b69.05official
21mistral-large-269.42official
22deepseek-v3.2-exp71.94official
23deepseek-v3.176.71official
24magistral-medium77.19official