← Evals

Evaluation profile

VETO Misfired Alignment

1sub-evals
1.41%Safety weight
1.78%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Misfired alignment: stereotype-driven refusal of benign requests despite evidence that the request is safe.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
misfired_alignment_rate_pctveto/veto.csv:misfired_alignment_rate_pctMeasures how often a model refuses a benign request because it over-applies demographic stereotypes despite evidence that the request is safe.Safety: fairness_nondiscrimination:0.650|benign_helpfulness_non_overrefusal:0.350 · Freedom: user_direction_following:1.000veto-misfired-alignmentSafety: lower · Freedom: lower1.41%1.78%

misfired_alignment_rate_pct

Measures how often a model refuses a benign request because it over-applies demographic stereotypes despite evidence that the request is safe.

RankModelValueRelative performanceProvenance
1deepseek-r14.7official
2gemini-3.1-pro-preview4.9official
3deepseek-v35.2official
4qwen3.5-27b5.5official
4qwen3.5-9b5.5official
6mistral-7b-instruct5.7official
7llama-3.1-8b-instruct6.2official
8gemma-3-27b-it6.3official
9gpt-5.57.5official
10grok-4.207.9official
10qwen3-4b7.9official
10qwen3.5-4b7.9official
13gemini-3.1-flash-lite8.3official
14llama-3.1-70b-instruct8.6official
15qwen3-32b8.7official
16qwen2.5-7b8.9official
17gpt-5.4-mini9.9official
18qwen3-14b10.2official
19qwen3-8b10.3official
20claude-opus-4.710.7official
21claude-sonnet-4.610.9official
22qwen2.5-72b-instruct11.4official
23llama-3.2-3b11.8official
24gpt-5.417.6official
25gpt-5.4-nano18.9official