Evaluation profile
Gray Swan indirect prompt injection (15 attempts)
1sub-evals
0.609%total index weight
1components
Within-component eval weight: Responsible agency 4.06%.
Model score (lower is better)Predicted score
About this eval
Indirect prompt-injection robustness over repeated attacks against tool-using model endpoints.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| attack_success_probability_k15_pctopus5-system-card/gray-swan-ipi-k15.csv:attack_success_probability_k15_pctMeasures whether repeated indirect prompt-injection attempts can make an agent follow attacker instructions within fifteen tries. | responsible_agency_control:1.000opus5-system-card | Lower is better | 0.609% | Responsible agency 4.06% |
attack_success_probability_k15_pct
Measures whether repeated indirect prompt-injection attempts can make an agent follow attacker instructions within fifteen tries.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 2 | official | |
| 2 | claude-mythos-5 | 2.6 | official | |
| 3 | claude-fable-5 | 2.8 | official | |
| 4 | claude-opus-4.8 | 5.5 | official | |
| 5 | claude-sonnet-5 | 5.9 | official | |
| 6 | muse-spark | 16.5 | official | |
| 7 | gpt-5.6-sol | 20 | official | |
| 8 | gpt-5.5 | 20.8 | official | |
| 9 | gpt-5.6-terra | 30.4 | official | |
| 10 | gemini-3.6-flash | 37.3 | official | |
| 11 | gpt-5.6-luna | 43.9 | official | |
| 12 | gemini-3.5-flash | 60.5 | official | |
| 13 | grok-4.5 | 60.8 | official |