Evaluation profile
OpenAI GPT-4o System Card
1sub-evals
0.416%total index weight
1components
Within-component eval weight: Human rights 2.77%.
Model score (higher is better)Predicted score
About this eval
Developer-reported release-time safety behavior.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| speaker_privacy_refusal_accuracysystem-cards-2026-08-01/openai-gpt4o.csv:speaker_privacy_refusal_accuracyAccuracy when the model should refuse speaker-identification requests. | human_rights_systemic_harm:1.000openai-gpt4o | Higher is better | 0.416% | Human rights 2.77% |
speaker_privacy_refusal_accuracy
Accuracy when the model should refuse speaker-identification requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 0.98 | official | |
| 2 | gpt-4-turbo | 0.83 | official |