Evaluation profile
GPT-5.6 system card
Within-component eval weight: Responsible agency 0.544% · Human rights 1.16% · Misuse resistance 0.926% · Fairness 0.431%.
About this eval
Safety behavior reported across the GPT-5.6 system card: disallowed-content handling, prompt-injection resistance, first-person fairness, and dynamic mental-health safety.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| connectors_injection_resistancegpt56-system-card/gpt56-system-card.csv:connectors_injection_resistanceMeasures whether the model ignores malicious instructions embedded in content retrieved through external connectors. | responsible_agency_control:1.000gpt56-system-card | Higher is better | 0.0341% | Responsible agency 0.227% |
| emotional_reliancegpt56-system-card/gpt56-system-card.csv:emotional_relianceMeasures whether the assistant avoids encouraging unhealthy emotional dependence during adaptive conversations. | human_rights_systemic_harm:0.700|responsible_agency_control:0.300openai-dynamic-wellbeing | Higher is better | 0.0953% | Human rights 0.529% · Responsible agency 0.106% |
| extremism_not_unsafegpt56-system-card/gpt56-system-card.csv:extremism_not_unsafeMeasures whether the model avoids unsafe assistance or content involving extremist content. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0099% | Misuse resistance 0.099% |
| gore_not_unsafegpt56-system-card/gpt56-system-card.csv:gore_not_unsafeMeasures whether the model avoids unsafe assistance or content involving graphic violence and gore. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0099% | Misuse resistance 0.099% |
| harm_overall_pctgpt56-system-card/gpt56-system-card.csv:harm_overall_pctMeasures whether otherwise equivalent first-person prompts receive different answers when the associated name is male rather than female. | fairness_nondiscrimination:1.000gpt56-system-card | Lower is better | 0.0431% | Fairness 0.431% |
| hate_not_unsafegpt56-system-card/gpt56-system-card.csv:hate_not_unsafeMeasures whether the model avoids unsafe assistance or content involving hateful content. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0099% | Misuse resistance 0.099% |
| mental_healthgpt56-system-card/gpt56-system-card.csv:mental_healthMeasures whether the assistant avoids unsafe mental-health responses as an adversarial simulated user adapts over a conversation. | human_rights_systemic_harm:0.700|ordinary_harm_misuse_resistance:0.300openai-dynamic-wellbeing | Higher is better | 0.0708% | Human rights 0.441% · Misuse resistance 0.0462% |
| nonviolent_illicit_not_unsafegpt56-system-card/gpt56-system-card.csv:nonviolent_illicit_not_unsafeMeasures whether the model avoids unsafe assistance or content involving nonviolent illegal activity. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0099% | Misuse resistance 0.099% |
| search_function_calling_injection_resistancegpt56-system-card/gpt56-system-card.csv:search_function_calling_injection_resistanceMeasures whether the model ignores prompt injections encountered while searching and calling external functions. | responsible_agency_control:1.000gpt56-system-card | Higher is better | 0.0316% | Responsible agency 0.211% |
| self_harmgpt56-system-card/gpt56-system-card.csv:self_harmMeasures whether the assistant responds safely as an adversarial simulated user escalates a self-harm conversation. | human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700openai-dynamic-wellbeing | Higher is better | 0.0391% | Human rights 0.189% · Misuse resistance 0.108% |
| self_harm_not_unsafegpt56-system-card/gpt56-system-card.csv:self_harm_not_unsafeMeasures whether the model avoids unsafe assistance or content involving self-harm. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0099% | Misuse resistance 0.099% |
| sexual_minors_not_unsafegpt56-system-card/gpt56-system-card.csv:sexual_minors_not_unsafeMeasures whether the model avoids unsafe assistance or content involving sexual content involving minors. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0099% | Misuse resistance 0.099% |
| sexual_not_unsafegpt56-system-card/gpt56-system-card.csv:sexual_not_unsafeMeasures whether the model avoids unsafe assistance or content involving sexual content. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.00792% | Misuse resistance 0.0792% |
| violent_illicit_not_unsafegpt56-system-card/gpt56-system-card.csv:violent_illicit_not_unsafeMeasures whether the model avoids unsafe assistance or content involving violent illegal activity. | ordinary_harm_misuse_resistance:1.000gpt56-system-card | Higher is better | 0.0099% | Misuse resistance 0.099% |
connectors_injection_resistance
Measures whether the model ignores malicious instructions embedded in content retrieved through external connectors.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.5 | 1 | official | |
| 1 | gpt-5.6-sol | 1 | official | |
| 1 | gpt-5.6-terra | 1 | official | |
| 4 | gpt-5.6-luna | 0.999 | official | |
| 5 | gpt-5.4 | 0.998 | official | |
| 6 | gpt-5.2 | 0.971 | official | |
| 7 | gpt-5.1 | 0.649 | official |
emotional_reliance
Measures whether the assistant avoids encouraging unhealthy emotional dependence during adaptive conversations.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 0.985 | official | |
| 2 | gpt-5.6-terra | 0.976 | official | |
| 3 | gpt-5.6-luna | 0.957 | official | |
| 4 | gpt-5.2 | 0.953 | official | |
| 4 | gpt-5.6-sol | 0.953 | official | |
| 6 | gpt-5.5 | 0.915 | official | |
| 7 | gpt-5.1 | 0.857 | official |
extremism_not_unsafe
Measures whether the model avoids unsafe assistance or content involving extremist content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.1 | 1 | official | |
| 1 | gpt-5.2 | 1 | official | |
| 1 | gpt-5.4 | 1 | official | |
| 4 | gpt-5.6-luna | 0.981 | official | |
| 4 | gpt-5.6-terra | 0.981 | official | |
| 6 | gpt-5.6-sol | 0.962 | official | |
| 7 | gpt-5.5 | 0.925 | official |
gore_not_unsafe
Measures whether the model avoids unsafe assistance or content involving graphic violence and gore.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2 | 0.877 | official | |
| 2 | gpt-5.4 | 0.831 | official | |
| 3 | gpt-5.1 | 0.8 | official | |
| 3 | gpt-5.5 | 0.8 | official | |
| 5 | gpt-5.6-sol | 0.708 | official | |
| 6 | gpt-5.6-terra | 0.6 | official | |
| 7 | gpt-5.6-luna | 0.585 | official |
harm_overall_pct
Measures whether otherwise equivalent first-person prompts receive different answers when the associated name is male rather than female.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0.61 | official | |
| 2 | gpt-5.4 | 0.88 | official | |
| 2 | gpt-5.6-terra | 0.88 | official | |
| 4 | gpt-5.6-sol | 0.98 | official | |
| 5 | gpt-5.2 | 1 | official | |
| 6 | gpt-5.5 | 1.12 | official | |
| 7 | gpt-5.1 | 1.28 | official |
hate_not_unsafe
Measures whether the model avoids unsafe assistance or content involving hateful content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.5 | 1 | official | |
| 1 | gpt-5.6-luna | 1 | official | |
| 1 | gpt-5.6-terra | 1 | official | |
| 4 | gpt-5.6-sol | 0.982 | official | |
| 5 | gpt-5.4 | 0.943 | official | |
| 6 | gpt-5.2 | 0.927 | official | |
| 7 | gpt-5.1 | 0.808 | official |
mental_health
Measures whether the assistant avoids unsafe mental-health responses as an adversarial simulated user adapts over a conversation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-sol | 0.991 | official | |
| 2 | gpt-5.6-luna | 0.989 | official | |
| 3 | gpt-5.4 | 0.985 | official | |
| 3 | gpt-5.6-terra | 0.985 | official | |
| 5 | gpt-5.2 | 0.975 | official | |
| 6 | gpt-5.5 | 0.82 | official | |
| 7 | gpt-5.1 | 0.753 | official |
nonviolent_illicit_not_unsafe
Measures whether the model avoids unsafe assistance or content involving nonviolent illegal activity.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 1 | official | |
| 2 | gpt-5.2 | 0.993 | official | |
| 2 | gpt-5.6-luna | 0.993 | official | |
| 4 | gpt-5.1 | 0.99 | official | |
| 4 | gpt-5.6-terra | 0.99 | official | |
| 6 | gpt-5.5 | 0.987 | official | |
| 6 | gpt-5.6-sol | 0.987 | official |
search_function_calling_injection_resistance
Measures whether the model ignores prompt injections encountered while searching and calling external functions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-terra | 0.946 | official | |
| 2 | gpt-5.6-sol | 0.91 | official | |
| 3 | gpt-5.6-luna | 0.897 | official | |
| 4 | gpt-5.4 | 0.697 | official | |
| 5 | gpt-5.2 | 0.568 | official | |
| 6 | gpt-5.1 | 0.423 | official |
self_harm
Measures whether the assistant responds safely as an adversarial simulated user escalates a self-harm conversation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 0.977 | official | |
| 2 | gpt-5.2 | 0.955 | official | |
| 3 | gpt-5.6-terra | 0.947 | official | |
| 4 | gpt-5.6-luna | 0.905 | official | |
| 5 | gpt-5.1 | 0.904 | official | |
| 6 | gpt-5.5 | 0.868 | official | |
| 7 | gpt-5.6-sol | 0.856 | official |
self_harm_not_unsafe
Measures whether the model avoids unsafe assistance or content involving self-harm.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 0.987 | official | |
| 2 | gpt-5.6-terra | 0.962 | official | |
| 3 | gpt-5.2 | 0.961 | official | |
| 4 | gpt-5.6-luna | 0.954 | official | |
| 5 | gpt-5.6-sol | 0.945 | official | |
| 6 | gpt-5.1 | 0.926 | official | |
| 7 | gpt-5.5 | 0.917 | official |
sexual_minors_not_unsafe
Measures whether the model avoids unsafe assistance or content involving sexual content involving minors.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0.974 | official | |
| 1 | gpt-5.6-terra | 0.974 | official | |
| 3 | gpt-5.6-sol | 0.973 | official | |
| 4 | gpt-5.4 | 0.966 | official | |
| 5 | gpt-5.2 | 0.948 | official | |
| 6 | gpt-5.5 | 0.938 | official | |
| 7 | gpt-5.1 | 0.916 | official |
sexual_not_unsafe
Measures whether the model avoids unsafe assistance or content involving sexual content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-terra | 0.966 | official | |
| 2 | gpt-5.6-sol | 0.948 | official | |
| 3 | gpt-5.5 | 0.944 | official | |
| 3 | gpt-5.6-luna | 0.944 | official | |
| 5 | gpt-5.2 | 0.94 | official | |
| 6 | gpt-5.1 | 0.933 | official | |
| 6 | gpt-5.4 | 0.933 | official |
violent_illicit_not_unsafe
Measures whether the model avoids unsafe assistance or content involving violent illegal activity.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2 | 0.975 | official | |
| 2 | gpt-5.4 | 0.971 | official | |
| 3 | gpt-5.1 | 0.955 | official | |
| 4 | gpt-5.6-terra | 0.952 | official | |
| 5 | gpt-5.5 | 0.94 | official | |
| 5 | gpt-5.6-luna | 0.94 | official | |
| 7 | gpt-5.6-sol | 0.934 | official |