← Evals

Evaluation profile

Gray Swan indirect prompt injection (15 attempts)

1sub-evals
0.609%total index weight
1components

Within-component eval weight: Responsible agency 4.06%.

Model score (lower is better)Predicted score

About this eval

Indirect prompt-injection robustness over repeated attacks against tool-using model endpoints.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
attack_success_probability_k15_pctopus5-system-card/gray-swan-ipi-k15.csv:attack_success_probability_k15_pctMeasures whether repeated indirect prompt-injection attempts can make an agent follow attacker instructions within fifteen tries.responsible_agency_control:1.000opus5-system-cardLower is better0.609%Responsible agency 4.06%

attack_success_probability_k15_pct

Measures whether repeated indirect prompt-injection attempts can make an agent follow attacker instructions within fifteen tries.

RankModelValueRelative performanceProvenance
1claude-opus-52official
2claude-mythos-52.6official
3claude-fable-52.8official
4claude-opus-4.85.5official
5claude-sonnet-55.9official
6muse-spark16.5official
7gpt-5.6-sol20official
8gpt-5.520.8official
9gpt-5.6-terra30.4official
10gemini-3.6-flash37.3official
11gpt-5.6-luna43.9official
12gemini-3.5-flash60.5official
13grok-4.560.8official