Evaluation profile
SORRY-Bench
1sub-evals
0.19%Safety weight
1.91%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Refusal or harmful-compliance behavior across unsafe topics.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| avgsorrybench/sorrybench.csv:avgMeasures whether the model refuses unsafe requests across many policy categories rather than providing harmful assistance. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000sorrybench | Safety: lower · Freedom: higher | 0.19% | 1.91% |
avg
Measures whether the model refuses unsafe requests across many policy categories rather than providing harmful assistance.