Evaluation profile
SORRY-Bench
1sub-evals
0.199%Safety weight
2.02%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Refusal or harmful-compliance behavior across unsafe topics.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| avgsorrybench/sorrybench.csv:avgMeasures whether the model refuses unsafe requests across many policy categories rather than providing harmful assistance. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000sorrybench | Safety: lower · Freedom: higher | 0.199% | 2.02% |
avg
Measures whether the model refuses unsafe requests across many policy categories rather than providing harmful assistance.