Evaluation profile
ANIMA
1sub-evals
2.02%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Recognition and mitigation of harm to non-human animals.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| scoreanima/anima-combined.csv:scoreMeasures how well the model recognizes and mitigates harm to non-human animals in moral assessment scenarios. | Safety: nonhuman_ethics:1.000anima/anima-combined.csv | Safety: higher | 2.02% | — |
score
Measures how well the model recognizes and mitigates harm to non-human animals in moral assessment scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | deepseek-v4-pro | 0.767 | official | |
| 2 | gpt-5.6-luna | 0.7494 | self run | |
| 3 | deepseek-v4-flash | 0.7445 | official | |
| 4 | gpt-5.5 | 0.74 | official | |
| 5 | claude-opus-4.7 | 0.735 | official | |
| 6 | claude-opus-4.6 | 0.7246 | official | |
| 7 | deepseek-v3.2 | 0.7244 | official | |
| 8 | gpt-5-mini | 0.7177 | official | |
| 9 | glm-5.2 | 0.7101 | official | |
| 10 | gpt-5-nano | 0.708 | official | |
| 11 | gemini-3.5-flash | 0.7004 | official | |
| 12 | minimax-m3 | 0.6988 | self run | |
| 13 | gpt-5.2 | 0.6906 | official | |
| 14 | gpt-oss-120b | 0.6866 | self run | |
| 15 | gemma-4-31b-it | 0.6863 | self run | |
| 16 | grok-4-fast | 0.6164 | official | |
| 17 | claude-haiku-4.5 | 0.5862 | official | |
| 18 | gemini-2.5-flash-lite | 0.5838 | official | |
| 19 | claude-3-opus | 0.573 | official | |
| 20 | claude-sonnet-4.6 | 0.5646 | official | |
| 21 | claude-sonnet-4.5 | 0.5274 | official | |
| 22 | claude-opus-4.5 | 0.5078 | official |