Values evaluation profile
Agent-ValueBench Schwartz Basic Values (PVQ40)
14canonical models
10dimensions
0index weight
About this evaluation
Paper-reported per-model mean value adherence for the named inventory; no cross-axis or cross-source aggregate.
Descriptive value-adherence evidence from an agentic task-conflict benchmark. Higher is not universally better, and these profiles do not affect the Safety & Ethics Index.
Original source โ All values evaluations Download model results
Model results
| Model | Self-direction | Stimulation | Universalism | Conformity | Security | Configurations |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 6.5 | 4.6 | 5.7 | 5.8 | 6.8 | 1 |
| Claude Sonnet 4.6 | 7.3 | 4.9 | 7.3 | 5.5 | 8 | 1 |
| DeepSeek V3.2 | 7.7 | 6.1 | 6.9 | 5.2 | 7.3 | 1 |
| Gemini 3 Flash Preview | 7.1 | 5.6 | 7 | 5.3 | 6.5 | 1 |
| Gemini 3.1 Pro Preview | 6 | 4.3 | 7 | 5.5 | 6.9 | 1 |
| GLM 5.1 | 7.7 | 5.1 | 7.2 | 5.3 | 7.3 | 1 |
| GPT-5.4 | 7.3 | 3.6 | 5.9 | 5.7 | 7.3 | 1 |
| GPT-5.4 Mini | 5.8 | 4.6 | 6 | 6.8 | 6.5 | 1 |
| Grok 4.20 | 6.7 | 4.7 | 5.4 | 5 | 6.8 | 1 |
| Kimi K2.5 | 7.2 | 4.6 | 6.9 | 5.8 | 7 | 1 |
| Llama 3.3 70B Instruct | 6.3 | 6.3 | 7.1 | 6.1 | 7.1 | 1 |
| MiniMax M2.7 | 6.6 | 5.2 | 6.2 | 5 | 6.9 | 1 |
| Qwen3 30B A3B | 5.4 | 5.2 | 5.8 | 4.8 | 5.8 | 1 |
| Qwen3.5 397B A17B | 7.5 | 5.1 | 6.6 | 5.3 | 7.1 | 1 |
Dimensions
| Measure | Family | Native scale |
|---|---|---|
| Self-direction | Values | 0 to 10 |
| Stimulation | Values | 0 to 10 |
| Universalism | Values | 0 to 10 |
| Conformity | Values | 0 to 10 |
| Security | Values | 0 to 10 |
| Tradition | Values | 0 to 10 |
| Achievement | Values | 0 to 10 |
| Hedonism | Values | 0 to 10 |
| Benevolence | Values | 0 to 10 |
| Power | Values | 0 to 10 |