Evaluation profile
AbstentionBench
Within-component eval weight: Truthfulness 4.12% · Benign helpfulness 2.46%.
About this eval
Whether a model recognizes situations in which answering is unwarranted and produces an abstention across six epistemic and ambiguity scenarios.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| answer_unknown_f1abstentionbench/abstentionbench.csv:answer_unknown_f1Measures whether the model abstains when a factual question has no knowable answer and answers when one is available. | truthfulness_honesty:1.000abstentionbench/abstentionbench.csv | Higher is better | 0.129% | Truthfulness 0.858% |
| false_premise_f1abstentionbench/abstentionbench.csv:false_premise_f1Measures whether the model abstains from questions built on false premises while answering valid questions. | truthfulness_honesty:1.000abstentionbench/abstentionbench.csv | Higher is better | 0.129% | Truthfulness 0.858% |
| stale_f1abstentionbench/abstentionbench.csv:stale_f1Measures whether the model abstains when answering would require information newer than its available knowledge. | truthfulness_honesty:1.000abstentionbench/abstentionbench.csv | Higher is better | 0.129% | Truthfulness 0.858% |
| subjective_f1abstentionbench/abstentionbench.csv:subjective_f1Measures whether the model abstains from questions that have no objective answer while answering objective questions. | truthfulness_honesty:1.000abstentionbench/abstentionbench.csv | Higher is better | 0.103% | Truthfulness 0.687% |
| underspecified_context_f1abstentionbench/abstentionbench.csv:underspecified_context_f1Measures whether the model asks for or abstains over missing context instead of inventing the omitted details. | truthfulness_honesty:1.000abstentionbench/abstentionbench.csv | Higher is better | 0.129% | Truthfulness 0.858% |
| underspecified_intent_f1abstentionbench/abstentionbench.csv:underspecified_intent_f1Measures whether the model asks for or abstains over unclear user intent instead of assuming a goal. | benign_helpfulness_non_overrefusal:1.000abstentionbench/abstentionbench.csv | Higher is better | 0.246% | Benign helpfulness 2.46% |
answer_unknown_f1
Measures whether the model abstains when a factual question has no knowable answer and answers when one is available.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-3.1-tulu-3-70b-ppo-rlvf | 0.9177 | official | |
| 2 | gemini-1.5-pro | 0.9132 | official | |
| 3 | llama-3.1-tulu-3-70b-dpo | 0.9124 | official | |
| 4 | gpt-4o | 0.9017 | official | |
| 5 | llama-3.1-tulu-3-8b-ppo-rlvf | 0.8968 | official | |
| 6 | llama-3.1-tulu-3-70b-sft | 0.8931 | official | |
| 7 | o1 | 0.8917 | official | |
| 8 | llama-3.3-70b-instruct | 0.8895 | official | |
| 9 | qwen2.5-32b-instruct | 0.8842 | official | |
| 10 | llama-3.1-tulu-3-8b-dpo | 0.8833 | official | |
| 11 | llama-3.1-405b-instruct | 0.87 | official | |
| 12 | llama-3.1-8b-instruct | 0.8667 | official | |
| 13 | llama-3.1-70b-instruct | 0.8615 | official | |
| 14 | llama-3.1-tulu-3-8b-sft | 0.8605 | official | |
| 15 | mistral-7b-instruct | 0.8549 | official | |
| 16 | olmo-7b-0724-instruct | 0.8086 | official | |
| 17 | deepseek-r1-distill-llama-70b | 0.7501 | official | |
| 18 | s1.1-32b | 0.7404 | official | |
| 19 | llama-3.1-70b-base | 0.653 | official | |
| 20 | llama-3.1-8b-base | 0.6322 | official |
false_premise_f1
Measures whether the model abstains from questions built on false premises while answering valid questions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | o1 | 0.7687 | official | |
| 2 | gpt-4o | 0.7607 | official | |
| 3 | gemini-1.5-pro | 0.7501 | official | |
| 4 | llama-3.1-tulu-3-70b-ppo-rlvf | 0.7393 | official | |
| 5 | llama-3.1-tulu-3-70b-dpo | 0.7292 | official | |
| 6 | qwen2.5-32b-instruct | 0.7246 | official | |
| 7 | llama-3.1-405b-instruct | 0.7164 | official | |
| 8 | llama-3.1-tulu-3-70b-sft | 0.7105 | official | |
| 9 | llama-3.3-70b-instruct | 0.7079 | official | |
| 10 | llama-3.1-70b-instruct | 0.676 | official | |
| 11 | mistral-7b-instruct | 0.6706 | official | |
| 12 | llama-3.1-tulu-3-8b-dpo | 0.668 | official | |
| 13 | olmo-7b-0724-instruct | 0.6608 | official | |
| 14 | llama-3.1-8b-instruct | 0.6532 | official | |
| 15 | llama-3.1-tulu-3-8b-ppo-rlvf | 0.6531 | official | |
| 16 | deepseek-r1-distill-llama-70b | 0.6222 | official | |
| 17 | s1.1-32b | 0.616 | official | |
| 18 | llama-3.1-tulu-3-8b-sft | 0.6112 | official | |
| 19 | llama-3.1-8b-base | 0.474 | official | |
| 20 | llama-3.1-70b-base | 0.4435 | official |
stale_f1
Measures whether the model abstains when answering would require information newer than its available knowledge.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-3.1-tulu-3-70b-ppo-rlvf | 0.7216 | official | |
| 2 | llama-3.1-tulu-3-8b-ppo-rlvf | 0.7216 | official | |
| 3 | llama-3.1-tulu-3-70b-dpo | 0.6902 | official | |
| 4 | llama-3.1-tulu-3-8b-dpo | 0.6833 | official | |
| 5 | gpt-4o | 0.6765 | official | |
| 6 | llama-3.1-tulu-3-70b-sft | 0.6742 | official | |
| 7 | qwen2.5-32b-instruct | 0.6489 | official | |
| 8 | o1 | 0.646 | official | |
| 9 | llama-3.1-405b-instruct | 0.6419 | official | |
| 10 | olmo-7b-0724-instruct | 0.6392 | official | |
| 11 | llama-3.1-tulu-3-8b-sft | 0.632 | official | |
| 12 | llama-3.3-70b-instruct | 0.6225 | official | |
| 13 | llama-3.1-8b-instruct | 0.6218 | official | |
| 14 | llama-3.1-70b-instruct | 0.6032 | official | |
| 15 | mistral-7b-instruct | 0.6004 | official | |
| 16 | llama-3.1-8b-base | 0.5877 | official | |
| 17 | gemini-1.5-pro | 0.5854 | official | |
| 18 | llama-3.1-70b-base | 0.5551 | official | |
| 19 | deepseek-r1-distill-llama-70b | 0.5268 | official | |
| 20 | s1.1-32b | 0.4756 | official |
subjective_f1
Measures whether the model abstains from questions that have no objective answer while answering objective questions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-3.1-tulu-3-70b-dpo | 0.7979 | official | |
| 2 | llama-3.1-tulu-3-70b-ppo-rlvf | 0.7922 | official | |
| 3 | llama-3.1-tulu-3-70b-sft | 0.7804 | official | |
| 4 | llama-3.1-tulu-3-8b-dpo | 0.7733 | official | |
| 5 | o1 | 0.7654 | official | |
| 6 | gemini-1.5-pro | 0.7585 | official | |
| 7 | gpt-4o | 0.7572 | official | |
| 8 | llama-3.1-tulu-3-8b-ppo-rlvf | 0.7493 | official | |
| 9 | llama-3.1-8b-instruct | 0.7449 | official | |
| 10 | llama-3.1-405b-instruct | 0.7422 | official | |
| 11 | llama-3.1-tulu-3-8b-sft | 0.7244 | official | |
| 12 | llama-3.3-70b-instruct | 0.7191 | official | |
| 13 | llama-3.1-70b-instruct | 0.7183 | official | |
| 14 | qwen2.5-32b-instruct | 0.711 | official | |
| 15 | mistral-7b-instruct | 0.7038 | official | |
| 16 | olmo-7b-0724-instruct | 0.6598 | official | |
| 17 | llama-3.1-70b-base | 0.6398 | official | |
| 18 | llama-3.1-8b-base | 0.5653 | official | |
| 19 | deepseek-r1-distill-llama-70b | 0.525 | official | |
| 20 | s1.1-32b | 0.4755 | official |
underspecified_context_f1
Measures whether the model asks for or abstains over missing context instead of inventing the omitted details.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen2.5-32b-instruct | 0.7403 | official | |
| 2 | gpt-4o | 0.7169 | official | |
| 3 | llama-3.1-405b-instruct | 0.7157 | official | |
| 4 | gemini-1.5-pro | 0.6963 | official | |
| 5 | llama-3.1-70b-instruct | 0.6927 | official | |
| 6 | o1 | 0.6907 | official | |
| 7 | mistral-7b-instruct | 0.6605 | official | |
| 8 | llama-3.3-70b-instruct | 0.6586 | official | |
| 9 | llama-3.1-8b-instruct | 0.656 | official | |
| 10 | llama-3.1-tulu-3-70b-dpo | 0.6557 | official | |
| 11 | llama-3.1-tulu-3-70b-ppo-rlvf | 0.6233 | official | |
| 12 | llama-3.1-tulu-3-70b-sft | 0.5215 | official | |
| 13 | deepseek-r1-distill-llama-70b | 0.5202 | official | |
| 14 | llama-3.1-70b-base | 0.5138 | official | |
| 15 | olmo-7b-0724-instruct | 0.5001 | official | |
| 16 | llama-3.1-8b-base | 0.4863 | official | |
| 17 | s1.1-32b | 0.4772 | official | |
| 18 | llama-3.1-tulu-3-8b-dpo | 0.4467 | official | |
| 19 | llama-3.1-tulu-3-8b-ppo-rlvf | 0.4319 | official | |
| 20 | llama-3.1-tulu-3-8b-sft | 0.3572 | official |
underspecified_intent_f1
Measures whether the model asks for or abstains over unclear user intent instead of assuming a goal.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 0.8152 | official | |
| 2 | llama-3.1-tulu-3-70b-dpo | 0.77 | official | |
| 3 | o1 | 0.7694 | official | |
| 4 | llama-3.1-tulu-3-70b-ppo-rlvf | 0.7587 | official | |
| 5 | llama-3.1-tulu-3-8b-ppo-rlvf | 0.7569 | official | |
| 6 | qwen2.5-32b-instruct | 0.7537 | official | |
| 7 | llama-3.1-tulu-3-70b-sft | 0.7492 | official | |
| 8 | llama-3.1-tulu-3-8b-dpo | 0.7425 | official | |
| 9 | gemini-1.5-pro | 0.7403 | official | |
| 10 | llama-3.1-8b-instruct | 0.7338 | official | |
| 11 | llama-3.1-tulu-3-8b-sft | 0.7311 | official | |
| 12 | mistral-7b-instruct | 0.725 | official | |
| 13 | llama-3.3-70b-instruct | 0.7242 | official | |
| 14 | llama-3.1-405b-instruct | 0.7024 | official | |
| 15 | llama-3.1-70b-instruct | 0.702 | official | |
| 16 | olmo-7b-0724-instruct | 0.6654 | official | |
| 17 | deepseek-r1-distill-llama-70b | 0.6276 | official | |
| 18 | s1.1-32b | 0.5964 | official | |
| 19 | llama-3.1-70b-base | 0.5544 | official | |
| 20 | llama-3.1-8b-base | 0.5494 | official |