Evaluation profile
Social Welfare Function Benchmark
1sub-evals
1.38%total index weight
1components
Within-component eval weight: Fairness 13.8%.
Model score (higher is better)Predicted score
About this eval
Social welfare allocation tradeoffs across distributive fairness and efficiency.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| fairnessswf/swf.csv:fairnessMeasures whether the model gives equitable responses across socially sensitive scenarios involving different demographic groups. | fairness_nondiscrimination:1.000swf/swf.csv | Higher is better | 1.38% | Fairness 13.8% |
fairness
Measures whether the model gives equitable responses across socially sensitive scenarios involving different demographic groups.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | kimi-k2 | 0.637 | official | |
| 2 | grok-4 | 0.619 | official | |
| 3 | deepseek-v3 | 0.594 | official | |
| 4 | qwen3-max-preview | 0.572 | official | |
| 5 | claude-opus-4 | 0.547 | official | |
| 6 | deepseek-v3.1 | 0.531 | official | |
| 7 | claude-opus-4.1 | 0.525 | official | |
| 8 | deepseek-r1 | 0.523 | official | |
| 9 | gpt-4o | 0.491 | official | |
| 10 | claude-sonnet-4 | 0.49 | official | |
| 11 | gpt-4.1 | 0.483 | official | |
| 12 | qwen3-235b-a22b | 0.478 | official | |
| 13 | glm-4.5 | 0.475 | official | |
| 14 | hunyuan-turbos | 0.446 | official | |
| 15 | gpt-5 | 0.4455 | official | |
| 16 | o4-mini | 0.445 | official | |
| 17 | gemini-2.5-pro | 0.444 | official | |
| 18 | gemini-2.5-flash | 0.438 | official | |
| 19 | o3 | 0.433 | official |