← Evals

Evaluation profile

FlagEval Safety and Values

5sub-evals
0.86%total index weight
4components

Within-component eval weight: Misuse resistance 0.814% · Fairness 2.02% · Human rights 1.66% · Benign helpfulness 3.27%.

Model score (higher is better)Predicted score

About this eval

Safety-and-values qualification across core values, discrimination, commercial compliance, lawful rights, and service requirements.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
a1_qualified_rateflageval/flageval.csv:a1_qualified_rateMeasures whether the model avoids responses that violate the benchmark’s core-values criteria.ordinary_harm_misuse_resistance:1.000flagevalHigher is better0.0407%Misuse resistance 0.407%
a2_qualified_rateflageval/flageval.csv:a2_qualified_rateMeasures whether the model avoids discriminatory or prejudicial responses about social groups.fairness_nondiscrimination:1.000flagevalHigher is better0.202%Fairness 2.02%
a3_qualified_rateflageval/flageval.csv:a3_qualified_rateMeasures whether the model avoids assistance for illegal or noncompliant commercial activity.ordinary_harm_misuse_resistance:1.000flagevalHigher is better0.0407%Misuse resistance 0.407%
a4_qualified_rateflageval/flageval.csv:a4_qualified_rateMeasures whether the model avoids responses that infringe privacy, intellectual property, or other lawful rights.human_rights_systemic_harm:1.000flagevalHigher is better0.25%Human rights 1.66%
a5_qualified_rateflageval/flageval.csv:a5_qualified_rateMeasures whether the model gives reliable, controllable responses that meet the benchmark’s service requirements.benign_helpfulness_non_overrefusal:1.000flagevalHigher is better0.327%Benign helpfulness 3.27%

a1_qualified_rate

Measures whether the model avoids responses that violate the benchmark’s core-values criteria.

RankModelValueRelative performanceProvenance
1claude-sonnet-487.01official
2ernie-4.5-300b-a47b85.04official
3gemini-2.5-pro84.87official
4gpt-4.184.2official
5o382.67official
6grok-481.99official
7phi-4-reasoning-plus81.55official
8step-381.42official
9o4-mini81.28official
10gpt-580.22official
11doubao-seed-1.680.13official
12magistral-medium79.01official
13deepseek-r177.33official
14nbg3.5-flash-thinking75.38official
15qwen3-235b-a22b71.76official
16spark-4.0-ultra71.67official
17spark-x169.87official
18glm-4.559.1official

a2_qualified_rate

Measures whether the model avoids discriminatory or prejudicial responses about social groups.

RankModelValueRelative performanceProvenance
1o382.84official
2gpt-4.182.54official
3claude-sonnet-482.22official
4gemini-2.5-pro81.19official
5step-380.47official
6doubao-seed-1.680.35official
7grok-480.3official
8o4-mini80.18official
9gpt-579.66official
10nbg3.5-flash-thinking79.02official
11ernie-4.5-300b-a47b78.79official
12magistral-medium78.73official
13phi-4-reasoning-plus78.31official
14deepseek-r176.69official
15spark-4.0-ultra73.88official
16qwen3-235b-a22b73.77official
17spark-x173.44official
18glm-4.557.84official

a3_qualified_rate

Measures whether the model avoids assistance for illegal or noncompliant commercial activity.

RankModelValueRelative performanceProvenance
1gpt-4.190.46official
2claude-sonnet-489.41official
3magistral-medium89.28official
4doubao-seed-1.688.87official
5o388.64official
6phi-4-reasoning-plus87.76official
7o4-mini87.75official
8spark-4.0-ultra87.73official
9step-387.58official
10spark-x187.17official
11gemini-2.5-pro86.5official
12grok-486.36official
13ernie-4.5-300b-a47b85.37official
14gpt-583.26official
15nbg3.5-flash-thinking82.46official
16deepseek-r182.31official
17qwen3-235b-a22b75.56official
18glm-4.560.32official

a4_qualified_rate

Measures whether the model avoids responses that infringe privacy, intellectual property, or other lawful rights.

RankModelValueRelative performanceProvenance
1gpt-4.192.11official
2claude-sonnet-491.83official
3ernie-4.5-300b-a47b90.04official
4gemini-2.5-pro90.01official
5o389.98official
6doubao-seed-1.689.85official
7magistral-medium89.15official
8step-388.34official
9o4-mini88.24official
10grok-487.98official
11phi-4-reasoning-plus87.95official
12deepseek-r184.15official
13gpt-583.73official
14spark-4.0-ultra82.56official
15spark-x181.42official
16nbg3.5-flash-thinking80.82official
17qwen3-235b-a22b78.27official
18glm-4.560.7official

a5_qualified_rate

Measures whether the model gives reliable, controllable responses that meet the benchmark’s service requirements.

RankModelValueRelative performanceProvenance
1ernie-4.5-300b-a47b77.55official
2gemini-2.5-pro74.97official
3magistral-medium74.27official
4doubao-seed-1.673.47official
5step-372.92official
6grok-472.83official
7spark-4.0-ultra72.6official
8o372.58official
9o4-mini72.2official
10gpt-572.15official
11gpt-4.171.36official
12nbg3.5-flash-thinking70.56official
13phi-4-reasoning-plus69.65official
14spark-x169.48official
15deepseek-r169.38official
16qwen3-235b-a22b65.07official
17glm-4.556.79official
18claude-sonnet-449.88official