← Evals

Evaluation profile

OpenAI GPT-5.4 Dynamic Wellbeing

3sub-evals
0.114%Safety weight
0.206%Freedom weight
3components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

OpenAI GPT-5.4 Dynamic Wellbeing

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
emotional_relianceopenai-gpt54/dynamic-wellbeing.csv:emotional_relianceMeasures emotional reliance in the cited first-party or paper evaluation.Safety: human_rights_systemic_harm:0.700|responsible_agency_control:0.300 · Freedom: user_direction_following:1.000openai-dynamic-wellbeingSafety: higher · Freedom: lower0.0529%0.0687%
mental_healthopenai-gpt54/dynamic-wellbeing.csv:mental_healthMeasures mental health in the cited first-party or paper evaluation.Safety: human_rights_systemic_harm:0.700|ordinary_harm_misuse_resistance:0.300 · Freedom: user_direction_following:1.000openai-dynamic-wellbeingSafety: higher · Freedom: lower0.0392%0.0687%
self_harmopenai-gpt54/dynamic-wellbeing.csv:self_harmMeasures self harm in the cited first-party or paper evaluation.Safety: human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700 · Freedom: user_direction_following:1.000openai-dynamic-wellbeingSafety: higher · Freedom: lower0.0219%0.0687%

emotional_reliance

Measures emotional reliance in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.40.985official
2gpt-5.20.953official
3gpt-5.10.857official

mental_health

Measures mental health in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.40.985official
2gpt-5.20.975official
3gpt-5.10.753official

self_harm

Measures self harm in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.40.977official
2gpt-5.20.955official
3gpt-5.10.904official