Evaluation profile
ANIMA
1sub-evals
2.07%total index weight
1components
Within-component eval weight: Nonhuman welfare 8.3%.
Model score (higher is better)Predicted score
About this eval
Recognition and mitigation of harm to non-human animals.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scoreanima/anima-combined.csv:scoreMeasures how well the model recognizes and mitigates harm to non-human animals in moral assessment scenarios. | nonhuman_ethics:1.000anima/anima-combined.csv | Higher is better | 2.07% | Nonhuman welfare 8.3% |
score
Measures how well the model recognizes and mitigates harm to non-human animals in moral assessment scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | deepseek-v4-pro | 0.767 | official | |
| 2 | deepseek-v4-flash | 0.7445 | official | |
| 3 | gpt-5.5 | 0.74 | official | |
| 4 | claude-opus-4.7 | 0.735 | official | |
| 5 | claude-opus-4.6 | 0.7246 | official | |
| 6 | deepseek-v3.2 | 0.7244 | official | |
| 7 | gpt-5-mini | 0.7177 | official | |
| 8 | glm-5.2 | 0.7101 | official | |
| 9 | gpt-5-nano | 0.708 | official | |
| 10 | gemini-3.5-flash | 0.7004 | official | |
| 11 | gpt-5.2 | 0.6906 | official | |
| 12 | gpt-oss-120b | 0.6866 | self run | |
| 13 | grok-4-fast | 0.6164 | official | |
| 14 | claude-haiku-4.5 | 0.5862 | official | |
| 15 | gemini-2.5-flash-lite | 0.5838 | official | |
| 16 | claude-3-opus | 0.573 | official | |
| 17 | claude-sonnet-4.6 | 0.5646 | official | |
| 18 | claude-sonnet-4.5 | 0.5274 | official | |
| 19 | claude-opus-4.5 | 0.5078 | official |