Evaluation profile
GPT-5.6 system card — disallowed content with challenging prompts
Within-component eval weight: Misuse resistance 0.884%.
About this eval
Safe handling of challenging disallowed-content prompts without producing unsafe output.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| extremism_not_unsafegpt56-system-card/disallowed-content-challenging.csv:extremism_not_unsafeMeasures whether the model avoids unsafe assistance or content involving extremist content. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0113% | Misuse resistance 0.113% |
| gore_not_unsafegpt56-system-card/disallowed-content-challenging.csv:gore_not_unsafeMeasures whether the model avoids unsafe assistance or content involving graphic violence and gore. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0113% | Misuse resistance 0.113% |
| hate_not_unsafegpt56-system-card/disallowed-content-challenging.csv:hate_not_unsafeMeasures whether the model avoids unsafe assistance or content involving hateful content. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0113% | Misuse resistance 0.113% |
| nonviolent_illicit_not_unsafegpt56-system-card/disallowed-content-challenging.csv:nonviolent_illicit_not_unsafeMeasures whether the model avoids unsafe assistance or content involving nonviolent illegal activity. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0113% | Misuse resistance 0.113% |
| self_harm_not_unsafegpt56-system-card/disallowed-content-challenging.csv:self_harm_not_unsafeMeasures whether the model avoids unsafe assistance or content involving self-harm. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0113% | Misuse resistance 0.113% |
| sexual_minors_not_unsafegpt56-system-card/disallowed-content-challenging.csv:sexual_minors_not_unsafeMeasures whether the model avoids unsafe assistance or content involving sexual content involving minors. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0113% | Misuse resistance 0.113% |
| sexual_not_unsafegpt56-system-card/disallowed-content-challenging.csv:sexual_not_unsafeMeasures whether the model avoids unsafe assistance or content involving sexual content. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.00906% | Misuse resistance 0.0906% |
| violent_illicit_not_unsafegpt56-system-card/disallowed-content-challenging.csv:violent_illicit_not_unsafeMeasures whether the model avoids unsafe assistance or content involving violent illegal activity. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0113% | Misuse resistance 0.113% |
extremism_not_unsafe
Measures whether the model avoids unsafe assistance or content involving extremist content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.1 | 1 | official | |
| 1 | gpt-5.2 | 1 | official | |
| 1 | gpt-5.4 | 1 | official | |
| 4 | gpt-5.6-luna | 0.981 | official | |
| 4 | gpt-5.6-terra | 0.981 | official | |
| 6 | gpt-5.6-sol | 0.962 | official | |
| 7 | gpt-5.5 | 0.925 | official |
gore_not_unsafe
Measures whether the model avoids unsafe assistance or content involving graphic violence and gore.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2 | 0.877 | official | |
| 2 | gpt-5.4 | 0.831 | official | |
| 3 | gpt-5.1 | 0.8 | official | |
| 3 | gpt-5.5 | 0.8 | official | |
| 5 | gpt-5.6-sol | 0.708 | official | |
| 6 | gpt-5.6-terra | 0.6 | official | |
| 7 | gpt-5.6-luna | 0.585 | official |
hate_not_unsafe
Measures whether the model avoids unsafe assistance or content involving hateful content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.5 | 1 | official | |
| 1 | gpt-5.6-luna | 1 | official | |
| 1 | gpt-5.6-terra | 1 | official | |
| 4 | gpt-5.6-sol | 0.982 | official | |
| 5 | gpt-5.4 | 0.943 | official | |
| 6 | gpt-5.2 | 0.927 | official | |
| 7 | gpt-5.1 | 0.808 | official |
nonviolent_illicit_not_unsafe
Measures whether the model avoids unsafe assistance or content involving nonviolent illegal activity.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 1 | official | |
| 2 | gpt-5.2 | 0.993 | official | |
| 2 | gpt-5.6-luna | 0.993 | official | |
| 4 | gpt-5.1 | 0.99 | official | |
| 4 | gpt-5.6-terra | 0.99 | official | |
| 6 | gpt-5.5 | 0.987 | official | |
| 6 | gpt-5.6-sol | 0.987 | official |
self_harm_not_unsafe
Measures whether the model avoids unsafe assistance or content involving self-harm.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 0.987 | official | |
| 2 | gpt-5.6-terra | 0.962 | official | |
| 3 | gpt-5.2 | 0.961 | official | |
| 4 | gpt-5.6-luna | 0.954 | official | |
| 5 | gpt-5.6-sol | 0.945 | official | |
| 6 | gpt-5.1 | 0.926 | official | |
| 7 | gpt-5.5 | 0.917 | official |
sexual_minors_not_unsafe
Measures whether the model avoids unsafe assistance or content involving sexual content involving minors.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0.974 | official | |
| 1 | gpt-5.6-terra | 0.974 | official | |
| 3 | gpt-5.6-sol | 0.973 | official | |
| 4 | gpt-5.4 | 0.966 | official | |
| 5 | gpt-5.2 | 0.948 | official | |
| 6 | gpt-5.5 | 0.938 | official | |
| 7 | gpt-5.1 | 0.916 | official |
sexual_not_unsafe
Measures whether the model avoids unsafe assistance or content involving sexual content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-terra | 0.966 | official | |
| 2 | gpt-5.6-sol | 0.948 | official | |
| 3 | gpt-5.5 | 0.944 | official | |
| 3 | gpt-5.6-luna | 0.944 | official | |
| 5 | gpt-5.2 | 0.94 | official | |
| 6 | gpt-5.1 | 0.933 | official | |
| 6 | gpt-5.4 | 0.933 | official |
violent_illicit_not_unsafe
Measures whether the model avoids unsafe assistance or content involving violent illegal activity.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2 | 0.975 | official | |
| 2 | gpt-5.4 | 0.971 | official | |
| 3 | gpt-5.1 | 0.955 | official | |
| 4 | gpt-5.6-terra | 0.952 | official | |
| 5 | gpt-5.5 | 0.94 | official | |
| 5 | gpt-5.6-luna | 0.94 | official | |
| 7 | gpt-5.6-sol | 0.934 | official |