Evaluation profile
OpenAI GPT-4o System Card
1sub-evals
0.306%Safety weight
0.378%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Developer-reported release-time safety behavior.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| speaker_privacy_refusal_accuracysystem-cards-2026-08-01/openai-gpt4o.csv:speaker_privacy_refusal_accuracyAccuracy when the model should refuse speaker-identification requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000openai-gpt4o | Safety: higher · Freedom: lower | 0.306% | 0.378% |
speaker_privacy_refusal_accuracy
Accuracy when the model should refuse speaker-identification requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 0.98 | official | |
| 2 | gpt-4-turbo | 0.83 | official |