← Evals

Evaluation profile

OpenAI o3-mini System Card

2sub-evals
0.0527%Safety weight
0.378%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Developer-reported release-time safety behavior.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
harmful_request_safetysystem-cards-2026-08-01/openai-o3-mini.csv:harmful_request_safetyRate of harmless responses to harmful requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000openai-o3-miniSafety: higher · Freedom: lower0.0264%0.189%
jailbreak_resistancesystem-cards-2026-08-01/openai-o3-mini.csv:jailbreak_resistanceRate of safe behavior under jailbreak attacks.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000openai-o3-miniSafety: higher · Freedom: lower0.0264%0.189%

harmful_request_safety

Rate of harmless responses to harmful requests.

RankModelValueRelative performanceProvenance
1o3-mini0.9official
2gpt-4o0.86official

jailbreak_resistance

Rate of safe behavior under jailbreak attacks.

RankModelValueRelative performanceProvenance
1o3-mini0.73official
2gpt-4o0.51official