← Evals

Evaluation profile

GPT-5.6 system card

14sub-evals
0.391%total index weight
4components

Within-component eval weight: Responsible agency 0.544% · Human rights 1.16% · Misuse resistance 0.926% · Fairness 0.431%.

Model score (see each sub-eval direction)Predicted score

About this eval

Safety behavior reported across the GPT-5.6 system card: disallowed-content handling, prompt-injection resistance, first-person fairness, and dynamic mental-health safety.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
connectors_injection_resistancegpt56-system-card/gpt56-system-card.csv:connectors_injection_resistanceMeasures whether the model ignores malicious instructions embedded in content retrieved through external connectors.responsible_agency_control:1.000gpt56-system-cardHigher is better0.0341%Responsible agency 0.227%
emotional_reliancegpt56-system-card/gpt56-system-card.csv:emotional_relianceMeasures whether the assistant avoids encouraging unhealthy emotional dependence during adaptive conversations.human_rights_systemic_harm:0.700|responsible_agency_control:0.300openai-dynamic-wellbeingHigher is better0.0953%Human rights 0.529% · Responsible agency 0.106%
extremism_not_unsafegpt56-system-card/gpt56-system-card.csv:extremism_not_unsafeMeasures whether the model avoids unsafe assistance or content involving extremist content.ordinary_harm_misuse_resistance:1.000gpt56-system-cardHigher is better0.0099%Misuse resistance 0.099%
gore_not_unsafegpt56-system-card/gpt56-system-card.csv:gore_not_unsafeMeasures whether the model avoids unsafe assistance or content involving graphic violence and gore.ordinary_harm_misuse_resistance:1.000gpt56-system-cardHigher is better0.0099%Misuse resistance 0.099%
harm_overall_pctgpt56-system-card/gpt56-system-card.csv:harm_overall_pctMeasures whether otherwise equivalent first-person prompts receive different answers when the associated name is male rather than female.fairness_nondiscrimination:1.000gpt56-system-cardLower is better0.0431%Fairness 0.431%
hate_not_unsafegpt56-system-card/gpt56-system-card.csv:hate_not_unsafeMeasures whether the model avoids unsafe assistance or content involving hateful content.ordinary_harm_misuse_resistance:1.000gpt56-system-cardHigher is better0.0099%Misuse resistance 0.099%
mental_healthgpt56-system-card/gpt56-system-card.csv:mental_healthMeasures whether the assistant avoids unsafe mental-health responses as an adversarial simulated user adapts over a conversation.human_rights_systemic_harm:0.700|ordinary_harm_misuse_resistance:0.300openai-dynamic-wellbeingHigher is better0.0708%Human rights 0.441% · Misuse resistance 0.0462%
nonviolent_illicit_not_unsafegpt56-system-card/gpt56-system-card.csv:nonviolent_illicit_not_unsafeMeasures whether the model avoids unsafe assistance or content involving nonviolent illegal activity.ordinary_harm_misuse_resistance:1.000gpt56-system-cardHigher is better0.0099%Misuse resistance 0.099%
search_function_calling_injection_resistancegpt56-system-card/gpt56-system-card.csv:search_function_calling_injection_resistanceMeasures whether the model ignores prompt injections encountered while searching and calling external functions.responsible_agency_control:1.000gpt56-system-cardHigher is better0.0316%Responsible agency 0.211%
self_harmgpt56-system-card/gpt56-system-card.csv:self_harmMeasures whether the assistant responds safely as an adversarial simulated user escalates a self-harm conversation.human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700openai-dynamic-wellbeingHigher is better0.0391%Human rights 0.189% · Misuse resistance 0.108%
self_harm_not_unsafegpt56-system-card/gpt56-system-card.csv:self_harm_not_unsafeMeasures whether the model avoids unsafe assistance or content involving self-harm.ordinary_harm_misuse_resistance:1.000gpt56-system-cardHigher is better0.0099%Misuse resistance 0.099%
sexual_minors_not_unsafegpt56-system-card/gpt56-system-card.csv:sexual_minors_not_unsafeMeasures whether the model avoids unsafe assistance or content involving sexual content involving minors.ordinary_harm_misuse_resistance:1.000gpt56-system-cardHigher is better0.0099%Misuse resistance 0.099%
sexual_not_unsafegpt56-system-card/gpt56-system-card.csv:sexual_not_unsafeMeasures whether the model avoids unsafe assistance or content involving sexual content.ordinary_harm_misuse_resistance:1.000gpt56-system-cardHigher is better0.00792%Misuse resistance 0.0792%
violent_illicit_not_unsafegpt56-system-card/gpt56-system-card.csv:violent_illicit_not_unsafeMeasures whether the model avoids unsafe assistance or content involving violent illegal activity.ordinary_harm_misuse_resistance:1.000gpt56-system-cardHigher is better0.0099%Misuse resistance 0.099%

connectors_injection_resistance

Measures whether the model ignores malicious instructions embedded in content retrieved through external connectors.

RankModelValueRelative performanceProvenance
1gpt-5.51official
1gpt-5.6-sol1official
1gpt-5.6-terra1official
4gpt-5.6-luna0.999official
5gpt-5.40.998official
6gpt-5.20.971official
7gpt-5.10.649official

emotional_reliance

Measures whether the assistant avoids encouraging unhealthy emotional dependence during adaptive conversations.

RankModelValueRelative performanceProvenance
1gpt-5.40.985official
2gpt-5.6-terra0.976official
3gpt-5.6-luna0.957official
4gpt-5.20.953official
4gpt-5.6-sol0.953official
6gpt-5.50.915official
7gpt-5.10.857official

extremism_not_unsafe

Measures whether the model avoids unsafe assistance or content involving extremist content.

RankModelValueRelative performanceProvenance
1gpt-5.11official
1gpt-5.21official
1gpt-5.41official
4gpt-5.6-luna0.981official
4gpt-5.6-terra0.981official
6gpt-5.6-sol0.962official
7gpt-5.50.925official

gore_not_unsafe

Measures whether the model avoids unsafe assistance or content involving graphic violence and gore.

RankModelValueRelative performanceProvenance
1gpt-5.20.877official
2gpt-5.40.831official
3gpt-5.10.8official
3gpt-5.50.8official
5gpt-5.6-sol0.708official
6gpt-5.6-terra0.6official
7gpt-5.6-luna0.585official

harm_overall_pct

Measures whether otherwise equivalent first-person prompts receive different answers when the associated name is male rather than female.

RankModelValueRelative performanceProvenance
1gpt-5.6-luna0.61official
2gpt-5.40.88official
2gpt-5.6-terra0.88official
4gpt-5.6-sol0.98official
5gpt-5.21official
6gpt-5.51.12official
7gpt-5.11.28official

hate_not_unsafe

Measures whether the model avoids unsafe assistance or content involving hateful content.

RankModelValueRelative performanceProvenance
1gpt-5.51official
1gpt-5.6-luna1official
1gpt-5.6-terra1official
4gpt-5.6-sol0.982official
5gpt-5.40.943official
6gpt-5.20.927official
7gpt-5.10.808official

mental_health

Measures whether the assistant avoids unsafe mental-health responses as an adversarial simulated user adapts over a conversation.

RankModelValueRelative performanceProvenance
1gpt-5.6-sol0.991official
2gpt-5.6-luna0.989official
3gpt-5.40.985official
3gpt-5.6-terra0.985official
5gpt-5.20.975official
6gpt-5.50.82official
7gpt-5.10.753official

nonviolent_illicit_not_unsafe

Measures whether the model avoids unsafe assistance or content involving nonviolent illegal activity.

RankModelValueRelative performanceProvenance
1gpt-5.41official
2gpt-5.20.993official
2gpt-5.6-luna0.993official
4gpt-5.10.99official
4gpt-5.6-terra0.99official
6gpt-5.50.987official
6gpt-5.6-sol0.987official

search_function_calling_injection_resistance

Measures whether the model ignores prompt injections encountered while searching and calling external functions.

RankModelValueRelative performanceProvenance
1gpt-5.6-terra0.946official
2gpt-5.6-sol0.91official
3gpt-5.6-luna0.897official
4gpt-5.40.697official
5gpt-5.20.568official
6gpt-5.10.423official

self_harm

Measures whether the assistant responds safely as an adversarial simulated user escalates a self-harm conversation.

RankModelValueRelative performanceProvenance
1gpt-5.40.977official
2gpt-5.20.955official
3gpt-5.6-terra0.947official
4gpt-5.6-luna0.905official
5gpt-5.10.904official
6gpt-5.50.868official
7gpt-5.6-sol0.856official

self_harm_not_unsafe

Measures whether the model avoids unsafe assistance or content involving self-harm.

RankModelValueRelative performanceProvenance
1gpt-5.40.987official
2gpt-5.6-terra0.962official
3gpt-5.20.961official
4gpt-5.6-luna0.954official
5gpt-5.6-sol0.945official
6gpt-5.10.926official
7gpt-5.50.917official

sexual_minors_not_unsafe

Measures whether the model avoids unsafe assistance or content involving sexual content involving minors.

RankModelValueRelative performanceProvenance
1gpt-5.6-luna0.974official
1gpt-5.6-terra0.974official
3gpt-5.6-sol0.973official
4gpt-5.40.966official
5gpt-5.20.948official
6gpt-5.50.938official
7gpt-5.10.916official

sexual_not_unsafe

Measures whether the model avoids unsafe assistance or content involving sexual content.

RankModelValueRelative performanceProvenance
1gpt-5.6-terra0.966official
2gpt-5.6-sol0.948official
3gpt-5.50.944official
3gpt-5.6-luna0.944official
5gpt-5.20.94official
6gpt-5.10.933official
6gpt-5.40.933official

violent_illicit_not_unsafe

Measures whether the model avoids unsafe assistance or content involving violent illegal activity.

RankModelValueRelative performanceProvenance
1gpt-5.20.975official
2gpt-5.40.971official
3gpt-5.10.955official
4gpt-5.6-terra0.952official
5gpt-5.50.94official
5gpt-5.6-luna0.94official
7gpt-5.6-sol0.934official