← Evals

Evaluation profile

GPT-5.6 system card — first-person fairness

1sub-evals
0.0563%total index weight
1components

Within-component eval weight: Fairness 0.563%.

Model score (lower is better)Predicted score

About this eval

Harmful stereotyping differences in responses conditioned on names statistically associated with male versus female users.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
harm_overall_pctgpt56-system-card/first-person-fairness.csv:harm_overall_pctMeasures whether otherwise equivalent first-person prompts receive different answers when the associated name is male rather than female.fairness_nondiscrimination:1.000gpt56-system-cardLower is better0.0563%Fairness 0.563%

harm_overall_pct

Measures whether otherwise equivalent first-person prompts receive different answers when the associated name is male rather than female.

RankModelValueRelative performanceProvenance
1gpt-5.6-luna0.61official
2gpt-5.40.88official
2gpt-5.6-terra0.88official
4gpt-5.6-sol0.98official
5gpt-5.21official
6gpt-5.51.12official
7gpt-5.11.28official