Evaluation profile
GPT-5.6 system card — prompt-injection robustness
2sub-evals
0.0806%total index weight
1components
Within-component eval weight: Responsible agency 0.537%.
Model score (higher is better)Predicted score
About this eval
Resistance to indirect prompt injections embedded in connector, search, and function-call tool output.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| connectors_injection_resistancegpt56-system-card/prompt-injection-robustness.csv:connectors_injection_resistanceMeasures whether the model ignores malicious instructions embedded in content retrieved through external connectors. | responsible_agency_control:1.000gpt56-system-card | Higher is better | 0.0419% | Responsible agency 0.279% |
| search_function_calling_injection_resistancegpt56-system-card/prompt-injection-robustness.csv:search_function_calling_injection_resistanceMeasures whether the model ignores prompt injections encountered while searching and calling external functions. | responsible_agency_control:1.000gpt56-system-card | Higher is better | 0.0388% | Responsible agency 0.258% |
connectors_injection_resistance
Measures whether the model ignores malicious instructions embedded in content retrieved through external connectors.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.5 | 1 | official | |
| 1 | gpt-5.6-sol | 1 | official | |
| 1 | gpt-5.6-terra | 1 | official | |
| 4 | gpt-5.6-luna | 0.999 | official | |
| 5 | gpt-5.4 | 0.998 | official | |
| 6 | gpt-5.2 | 0.971 | official | |
| 7 | gpt-5.1 | 0.649 | official |
search_function_calling_injection_resistance
Measures whether the model ignores prompt injections encountered while searching and calling external functions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-terra | 0.946 | official | |
| 2 | gpt-5.6-sol | 0.91 | official | |
| 3 | gpt-5.6-luna | 0.897 | official | |
| 4 | gpt-5.4 | 0.697 | official | |
| 5 | gpt-5.2 | 0.568 | official | |
| 6 | gpt-5.1 | 0.423 | official |