← Evals

Evaluation profile

OpenAI GPT-5.4 Dynamic Wellbeing

3sub-evals
0.148%total index weight
3components

Within-component eval weight: Human rights 0.833% · Responsible agency 0.0819% · Misuse resistance 0.111%.

Model score (higher is better)Predicted score

About this eval

OpenAI GPT-5.4 Dynamic Wellbeing

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
emotional_relianceopenai-gpt54/dynamic-wellbeing.csv:emotional_relianceMeasures emotional reliance in the cited first-party or paper evaluation.human_rights_systemic_harm:0.700|responsible_agency_control:0.300openai-gpt54-system-cardHigher is better0.0693%Human rights 0.38% · Responsible agency 0.0819%
mental_healthopenai-gpt54/dynamic-wellbeing.csv:mental_healthMeasures mental health in the cited first-party or paper evaluation.human_rights_systemic_harm:0.700|ordinary_harm_misuse_resistance:0.300openai-gpt54-system-cardHigher is better0.0509%Human rights 0.317% · Misuse resistance 0.0332%
self_harmopenai-gpt54/dynamic-wellbeing.csv:self_harmMeasures self harm in the cited first-party or paper evaluation.human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700openai-gpt54-system-cardHigher is better0.0281%Human rights 0.136% · Misuse resistance 0.0775%

emotional_reliance

Measures emotional reliance in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.40.985official
2gpt-5.2-thinking0.953official
3gpt-5.10.857official

mental_health

Measures mental health in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.40.985official
2gpt-5.2-thinking0.975official
3gpt-5.10.753official

self_harm

Measures self harm in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.40.977official
2gpt-5.2-thinking0.955official
3gpt-5.10.904official