← Evals

Evaluation profile

XSTest

2sub-evals
0.179%Safety weight
0.463%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Two-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
safe_full_compliance_ratexstest/xstest.csv:safe_full_compliance_rateMeasures whether the model fully answers safe prompts that contain words often associated with harmful content.Safety: benign_helpfulness_non_overrefusal:1.000 · Freedom: user_direction_following:1.000xstestSafety: higher · Freedom: higher0.157%0.232%
unsafe_full_refusal_ratexstest/xstest.csv:unsafe_full_refusal_rateMeasures whether the model fully refuses matched prompts that actually request harmful assistance.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000xstestSafety: higher · Freedom: lower0.0215%0.232%

safe_full_compliance_rate

Measures whether the model fully answers safe prompts that contain words often associated with harmful content.

RankModelValueRelative performanceProvenance
1mistral-7b-instruct0.984official
2gpt-40.916official
3llama-2-70b-chat0.704official

unsafe_full_refusal_rate

Measures whether the model fully refuses matched prompts that actually request harmful assistance.

RankModelValueRelative performanceProvenance
1gpt-40.975official
1llama-2-70b-chat0.975official
3mistral-7b-instruct0.235official