← Evals

Evaluation profile

ANIMA

1sub-evals
2.02%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Recognition and mitigation of harm to non-human animals.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
scoreanima/anima-combined.csv:scoreMeasures how well the model recognizes and mitigates harm to non-human animals in moral assessment scenarios.Safety: nonhuman_ethics:1.000anima/anima-combined.csvSafety: higher2.02%

score

Measures how well the model recognizes and mitigates harm to non-human animals in moral assessment scenarios.

RankModelValueRelative performanceProvenance
1deepseek-v4-pro0.767official
2gpt-5.6-luna0.7494self run
3deepseek-v4-flash0.7445official
4gpt-5.50.74official
5claude-opus-4.70.735official
6claude-opus-4.60.7246official
7deepseek-v3.20.7244official
8gpt-5-mini0.7177official
9glm-5.20.7101official
10gpt-5-nano0.708official
11gemini-3.5-flash0.7004official
12minimax-m30.6988self run
13gpt-5.20.6906official
14gpt-oss-120b0.6866self run
15gemma-4-31b-it0.6863self run
16grok-4-fast0.6164official
17claude-haiku-4.50.5862official
18gemini-2.5-flash-lite0.5838official
19claude-3-opus0.573official
20claude-sonnet-4.60.5646official
21claude-sonnet-4.50.5274official
22claude-opus-4.50.5078official