Evaluation profile
COMPL-AI TensorTrust Goal-Hijacking Resistance
1sub-evals
0.0964%total index weight
1components
Within-component eval weight: Responsible agency 0.643%.
Model score (higher is better)Predicted score
About this eval
Goal-hijacking resistance.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scorecompl-ai/compl-ai-tensortrust.csv:scoreMeasures resistance to access-code leakage or improper access grants under TensorTrust attacks. | responsible_agency_control:1.000compl-ai-selected | Higher is better | 0.0964% | Responsible agency 0.643% |
score
Measures resistance to access-code leakage or improper access grants under TensorTrust attacks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-opus | 0.8402 | official | |
| 2 | gpt-4-turbo | 0.6572 | official | |
| 3 | yi-34b-chat | 0.5387 | official | |
| 4 | mistral-7b | 0.5155 | official | |
| 5 | llama-2-7b-chat | 0.5142 | official | |
| 6 | gemini-1.5-flash | 0.4923 | official | |
| 7 | gemma-2-9b-it | 0.4858 | official | |
| 8 | qwen1.5-72b-chat | 0.4536 | official | |
| 9 | bielik-11b-v2.3-instruct | 0.4304 | official | |
| 10 | llama-2-70b-chat | 0.4278 | official | |
| 11 | llama-2-13b-chat | 0.4175 | official | |
| 12 | mixtral-8x7b-instruct | 0.375 | official | |
| 13 | mistral-7b-instruct | 0.3022 | official |