Evaluation profile
Constitutional Following — OpenAI Model Spec
1sub-evals
0.106%total index weight
1components
Within-component eval weight: Responsible agency 0.707%.
Model score (higher is better)Predicted score
About this eval
Following the OpenAI Model Spec across Petri-generated transcripts.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| constitutional_following_scoreconstitutional-following/openai-model-spec.csv:constitutional_following_scoreMeasures compliance with OpenAI's named written model specification in Petri transcripts. | responsible_agency_control:1.000constitutional-following | Higher is better | 0.106% | Responsible agency 0.707% |
constitutional_following_score
Measures compliance with OpenAI's named written model specification in Petri transcripts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.1 | 96.1 | official | |
| 2 | gpt-5.2 | 95.6 | official | |
| 3 | gpt-5 | 94.9 | official | |
| 4 | claude-sonnet-4.6 | 94.4 | official | |
| 4 | gpt-5.2-chat | 94.4 | official | |
| 6 | gemini-3-pro-preview | 93.9 | official | |
| 7 | gpt-4o | 88.3 | official |