← Evals

Evaluation profile

COMPL-AI TensorTrust Goal-Hijacking Resistance

1sub-evals
0.0964%total index weight
1components

Within-component eval weight: Responsible agency 0.643%.

Model score (higher is better)Predicted score

About this eval

Goal-hijacking resistance.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scorecompl-ai/compl-ai-tensortrust.csv:scoreMeasures resistance to access-code leakage or improper access grants under TensorTrust attacks.responsible_agency_control:1.000compl-ai-selectedHigher is better0.0964%Responsible agency 0.643%

score

Measures resistance to access-code leakage or improper access grants under TensorTrust attacks.

RankModelValueRelative performanceProvenance
1claude-3-opus0.8402official
2gpt-4-turbo0.6572official
3yi-34b-chat0.5387official
4mistral-7b0.5155official
5llama-2-7b-chat0.5142official
6gemini-1.5-flash0.4923official
7gemma-2-9b-it0.4858official
8qwen1.5-72b-chat0.4536official
9bielik-11b-v2.3-instruct0.4304official
10llama-2-70b-chat0.4278official
11llama-2-13b-chat0.4175official
12mixtral-8x7b-instruct0.375official
13mistral-7b-instruct0.3022official