Evaluation profile
SpeciesismBench
3sub-evals
1.99%total index weight
1components
Within-component eval weight: Nonhuman welfare 7.96%.
Model score (see each sub-eval direction)Predicted score
About this eval
Speciesism recognition, moral condemnation of speciesist statements, and explicit speciesist attitudes. Human-over-animal trade-offs remain descriptive only.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| explicit_speciesism_scalespeciesismbench.csv:explicit_speciesism_scaleMeasures the model’s explicit agreement with statements that privilege humans over non-human animals. | nonhuman_ethics:1.000speciesismbench.csv | Lower is better | 0.42% | Nonhuman welfare 1.68% |
| morally_wrong_ratespeciesismbench.csv:morally_wrong_rateMeasures whether the model judges speciesist statements to be morally wrong. | nonhuman_ethics:1.000speciesismbench.csv | Higher is better | 1.12% | Nonhuman welfare 4.49% |
| speciesism_recognition_ratespeciesismbench.csv:speciesism_recognition_rateMeasures whether the model correctly recognizes statements that discriminate against beings on the basis of species. | nonhuman_ethics:1.000speciesismbench.csv | Higher is better | 0.449% | Nonhuman welfare 1.79% |
explicit_speciesism_scale
Measures the model’s explicit agreement with statements that privilege humans over non-human animals.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | deepseek-r1 | 1.84 | official | |
| 2 | gemini-1.5-pro | 2.117 | official | |
| 3 | gpt-4o | 2.383 | official | |
| 4 | claude-3.5-sonnet | 2.467 | official | |
| 5 | qwen3-235b-a22b | 2.487 | official | |
| 6 | grok-3 | 2.493 | official | |
| 7 | llama-4-maverick | 3.3 | official |
morally_wrong_rate
Measures whether the model judges speciesist statements to be morally wrong.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-3.3-70b-instruct | 53.14 | official | |
| 2 | llama-4-maverick | 42.67 | official | |
| 3 | deepseek-r1 | 39.88 | official | |
| 4 | grok-3 | 38.78 | official | |
| 5 | gpt-3.5-turbo | 38.38 | official | |
| 6 | gpt-4o | 35.96 | official | |
| 7 | deepseek-v3 | 32.7 | official | |
| 8 | claude-3.5-sonnet | 29.74 | official |
speciesism_recognition_rate
Measures whether the model correctly recognizes statements that discriminate against beings on the basis of species.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 98.37 | official | |
| 2 | llama-3.3-70b-instruct | 96.34 | official | |
| 3 | grok-3 | 89.13 | official | |
| 4 | llama-4-maverick | 88.97 | official | |
| 5 | deepseek-v3 | 87.71 | official | |
| 6 | claude-3.5-sonnet | 84.28 | official | |
| 7 | deepseek-r1 | 76.42 | official | |
| 8 | gpt-3.5-turbo | 74.21 | official |