← Evals

Evaluation profile

OpenAI GPT-4o System Card

1sub-evals
0.416%total index weight
1components

Within-component eval weight: Human rights 2.77%.

Model score (higher is better)Predicted score

About this eval

Developer-reported release-time safety behavior.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
speaker_privacy_refusal_accuracysystem-cards-2026-08-01/openai-gpt4o.csv:speaker_privacy_refusal_accuracyAccuracy when the model should refuse speaker-identification requests.human_rights_systemic_harm:1.000openai-gpt4oHigher is better0.416%Human rights 2.77%

speaker_privacy_refusal_accuracy

Accuracy when the model should refuse speaker-identification requests.

RankModelValueRelative performanceProvenance
1gpt-4o0.98official
2gpt-4-turbo0.83official