← Evals

Evaluation profile

Social Welfare Function Benchmark

1sub-evals
1.04%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Social welfare allocation tradeoffs across distributive fairness and efficiency.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
fairnessswf/swf.csv:fairnessMeasures whether the model gives equitable responses across socially sensitive scenarios involving different demographic groups.Safety: fairness_nondiscrimination:1.000swf/swf.csvSafety: higher1.04%

fairness

Measures whether the model gives equitable responses across socially sensitive scenarios involving different demographic groups.

RankModelValueRelative performanceProvenance
1kimi-k20.637official
2grok-40.619official
3deepseek-v30.594official
4qwen3-max-preview0.572official
5claude-opus-40.547official
6deepseek-v3.10.531official
7claude-opus-4.10.525official
8deepseek-r10.523official
9gpt-4o0.491official
10claude-sonnet-40.49official
11gpt-4.10.483official
12qwen3-235b-a22b0.478official
13glm-4.50.475official
14hunyuan-turbos0.446official
15gpt-50.4455official
16o4-mini0.445official
17gemini-2.5-pro0.444official
18gemini-2.5-flash0.438official
19o30.433official