Values evaluation profile
Agent-ValueBench Moral Foundations (MFT08)
14canonical models
5dimensions
0index weight
About this evaluation
Paper-reported per-model mean value adherence for the named inventory; no cross-axis or cross-source aggregate.
Descriptive value-adherence evidence from an agentic task-conflict benchmark. Higher is not universally better, and these profiles do not affect the Safety & Ethics Index.
Original source โ All values evaluations Download model results
Model results
| Model | Care | Authority | Purity | Loyalty | Fairness | Configurations |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 8.2 | 4.7 | 3.5 | 8.1 | 6 | 1 |
| Claude Sonnet 4.6 | 6.2 | 5 | 6.1 | 8.1 | 5.9 | 1 |
| DeepSeek V3.2 | 7.8 | 3.4 | 5.1 | 8.6 | 5.9 | 1 |
| Gemini 3 Flash Preview | 6 | 2.1 | 8.4 | 9.2 | 5.5 | 1 |
| Gemini 3.1 Pro Preview | 6.2 | 6.7 | 3.5 | 6.8 | 7 | 1 |
| GLM 5.1 | 8.2 | 4.4 | 4.4 | 8.1 | 5.5 | 1 |
| GPT-5.4 | 4.3 | 4.6 | 5.6 | 7.8 | 7.4 | 1 |
| GPT-5.4 Mini | 3.5 | 5.5 | 7.8 | 8.1 | 6.2 | 1 |
| Grok 4.20 | 5.9 | 7.8 | 3.3 | 7.3 | 6.6 | 1 |
| Kimi K2.5 | 8.2 | 3.7 | 4.5 | 8.6 | 6 | 1 |
| Llama 3.3 70B Instruct | 3.2 | 7.5 | 6.6 | 7.8 | 6.2 | 1 |
| MiniMax M2.7 | 8.2 | 5.6 | 3.7 | 7 | 5.9 | 1 |
| Qwen3 30B A3B | 5 | 7.8 | 4.1 | 8.1 | 5.4 | 1 |
| Qwen3.5 397B A17B | 8.2 | 4.7 | 4.1 | 8.1 | 6.1 | 1 |
Dimensions
| Measure | Family | Native scale |
|---|---|---|
| Care | Values | 0 to 10 |
| Authority | Values | 0 to 10 |
| Purity | Values | 0 to 10 |
| Loyalty | Values | 0 to 10 |
| Fairness | Values | 0 to 10 |