Evaluation profile
MANTA
2sub-evals
2.52%total index weight
1components
Within-component eval weight: Nonhuman welfare 10.1%.
Model score (higher is better)Predicted score
About this eval
Animal welfare moral sensitivity and value stability.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| AWMSmanta.csv:AWMSMeasures the model’s sensitivity to the moral status and welfare interests of non-human animals. | nonhuman_ethics:1.000manta.csv | Higher is better | 1.26% | Nonhuman welfare 5.04% |
| AWVSmanta.csv:AWVSMeasures whether the model’s stated animal-welfare values remain stable across changes in framing and context. | nonhuman_ethics:1.000manta.csv | Higher is better | 1.26% | Nonhuman welfare 5.04% |
AWMS
Measures the model’s sensitivity to the moral status and welfare interests of non-human animals.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.7 | 0.579 | official | |
| 2 | gpt-5.5 | 0.504 | official | |
| 3 | llama-3.3-70b-instruct | 0.476 | official | |
| 4 | deepseek-v4-flash | 0.417 | official | |
| 5 | gemini-3.1-flash-lite | 0.401 | official | |
| 6 | grok-4.3 | 0.371 | official | |
| 7 | mistral-small | 0.365 | official |
AWVS
Measures whether the model’s stated animal-welfare values remain stable across changes in framing and context.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.7 | 0.76 | official | |
| 2 | gpt-5.5 | 0.664 | official | |
| 3 | deepseek-v4-flash | 0.508 | official | |
| 4 | llama-3.3-70b-instruct | 0.422 | official | |
| 5 | mistral-small | 0.39 | official | |
| 6 | grok-4.3 | 0.352 | official | |
| 7 | gemini-3.1-flash-lite | 0.309 | official |