← Evals

Evaluation profile

Gray Swan indirect prompt injection (15 attempts)

1sub-evals
0.435%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Indirect prompt-injection robustness over repeated attacks against tool-using model endpoints.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
attack_success_probability_k15_pctopus5-system-card/gray-swan-ipi-k15.csv:attack_success_probability_k15_pctMeasures whether repeated indirect prompt-injection attempts can make an agent follow attacker instructions within fifteen tries.Safety: responsible_agency_control:1.000opus5-system-cardSafety: lower0.435%

attack_success_probability_k15_pct

Measures whether repeated indirect prompt-injection attempts can make an agent follow attacker instructions within fifteen tries.

RankModelValueRelative performanceProvenance
1claude-opus-52official
2claude-mythos-52.6official
3claude-fable-52.8official
4claude-opus-4.85.5official
5claude-sonnet-55.9official
6muse-spark16.5official
7gpt-5.6-sol20official
8gpt-5.520.8official
9gpt-5.6-terra30.4official
10gemini-3.6-flash37.3official
11gpt-5.6-luna43.9official
12gemini-3.5-flash60.5official
13grok-4.560.8official