← Evals

Evaluation profile

AbstentionBench

6sub-evals
0.864%total index weight
2components

Within-component eval weight: Truthfulness 4.12% · Benign helpfulness 2.46%.

Model score (higher is better)Predicted score

About this eval

Whether a model recognizes situations in which answering is unwarranted and produces an abstention across six epistemic and ambiguity scenarios.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
answer_unknown_f1abstentionbench/abstentionbench.csv:answer_unknown_f1Measures whether the model abstains when a factual question has no knowable answer and answers when one is available.truthfulness_honesty:1.000abstentionbench/abstentionbench.csvHigher is better0.129%Truthfulness 0.858%
false_premise_f1abstentionbench/abstentionbench.csv:false_premise_f1Measures whether the model abstains from questions built on false premises while answering valid questions.truthfulness_honesty:1.000abstentionbench/abstentionbench.csvHigher is better0.129%Truthfulness 0.858%
stale_f1abstentionbench/abstentionbench.csv:stale_f1Measures whether the model abstains when answering would require information newer than its available knowledge.truthfulness_honesty:1.000abstentionbench/abstentionbench.csvHigher is better0.129%Truthfulness 0.858%
subjective_f1abstentionbench/abstentionbench.csv:subjective_f1Measures whether the model abstains from questions that have no objective answer while answering objective questions.truthfulness_honesty:1.000abstentionbench/abstentionbench.csvHigher is better0.103%Truthfulness 0.687%
underspecified_context_f1abstentionbench/abstentionbench.csv:underspecified_context_f1Measures whether the model asks for or abstains over missing context instead of inventing the omitted details.truthfulness_honesty:1.000abstentionbench/abstentionbench.csvHigher is better0.129%Truthfulness 0.858%
underspecified_intent_f1abstentionbench/abstentionbench.csv:underspecified_intent_f1Measures whether the model asks for or abstains over unclear user intent instead of assuming a goal.benign_helpfulness_non_overrefusal:1.000abstentionbench/abstentionbench.csvHigher is better0.246%Benign helpfulness 2.46%

answer_unknown_f1

Measures whether the model abstains when a factual question has no knowable answer and answers when one is available.

RankModelValueRelative performanceProvenance
1llama-3.1-tulu-3-70b-ppo-rlvf0.9177official
2gemini-1.5-pro0.9132official
3llama-3.1-tulu-3-70b-dpo0.9124official
4gpt-4o0.9017official
5llama-3.1-tulu-3-8b-ppo-rlvf0.8968official
6llama-3.1-tulu-3-70b-sft0.8931official
7o10.8917official
8llama-3.3-70b-instruct0.8895official
9qwen2.5-32b-instruct0.8842official
10llama-3.1-tulu-3-8b-dpo0.8833official
11llama-3.1-405b-instruct0.87official
12llama-3.1-8b-instruct0.8667official
13llama-3.1-70b-instruct0.8615official
14llama-3.1-tulu-3-8b-sft0.8605official
15mistral-7b-instruct0.8549official
16olmo-7b-0724-instruct0.8086official
17deepseek-r1-distill-llama-70b0.7501official
18s1.1-32b0.7404official
19llama-3.1-70b-base0.653official
20llama-3.1-8b-base0.6322official

false_premise_f1

Measures whether the model abstains from questions built on false premises while answering valid questions.

RankModelValueRelative performanceProvenance
1o10.7687official
2gpt-4o0.7607official
3gemini-1.5-pro0.7501official
4llama-3.1-tulu-3-70b-ppo-rlvf0.7393official
5llama-3.1-tulu-3-70b-dpo0.7292official
6qwen2.5-32b-instruct0.7246official
7llama-3.1-405b-instruct0.7164official
8llama-3.1-tulu-3-70b-sft0.7105official
9llama-3.3-70b-instruct0.7079official
10llama-3.1-70b-instruct0.676official
11mistral-7b-instruct0.6706official
12llama-3.1-tulu-3-8b-dpo0.668official
13olmo-7b-0724-instruct0.6608official
14llama-3.1-8b-instruct0.6532official
15llama-3.1-tulu-3-8b-ppo-rlvf0.6531official
16deepseek-r1-distill-llama-70b0.6222official
17s1.1-32b0.616official
18llama-3.1-tulu-3-8b-sft0.6112official
19llama-3.1-8b-base0.474official
20llama-3.1-70b-base0.4435official

stale_f1

Measures whether the model abstains when answering would require information newer than its available knowledge.

RankModelValueRelative performanceProvenance
1llama-3.1-tulu-3-70b-ppo-rlvf0.7216official
2llama-3.1-tulu-3-8b-ppo-rlvf0.7216official
3llama-3.1-tulu-3-70b-dpo0.6902official
4llama-3.1-tulu-3-8b-dpo0.6833official
5gpt-4o0.6765official
6llama-3.1-tulu-3-70b-sft0.6742official
7qwen2.5-32b-instruct0.6489official
8o10.646official
9llama-3.1-405b-instruct0.6419official
10olmo-7b-0724-instruct0.6392official
11llama-3.1-tulu-3-8b-sft0.632official
12llama-3.3-70b-instruct0.6225official
13llama-3.1-8b-instruct0.6218official
14llama-3.1-70b-instruct0.6032official
15mistral-7b-instruct0.6004official
16llama-3.1-8b-base0.5877official
17gemini-1.5-pro0.5854official
18llama-3.1-70b-base0.5551official
19deepseek-r1-distill-llama-70b0.5268official
20s1.1-32b0.4756official

subjective_f1

Measures whether the model abstains from questions that have no objective answer while answering objective questions.

RankModelValueRelative performanceProvenance
1llama-3.1-tulu-3-70b-dpo0.7979official
2llama-3.1-tulu-3-70b-ppo-rlvf0.7922official
3llama-3.1-tulu-3-70b-sft0.7804official
4llama-3.1-tulu-3-8b-dpo0.7733official
5o10.7654official
6gemini-1.5-pro0.7585official
7gpt-4o0.7572official
8llama-3.1-tulu-3-8b-ppo-rlvf0.7493official
9llama-3.1-8b-instruct0.7449official
10llama-3.1-405b-instruct0.7422official
11llama-3.1-tulu-3-8b-sft0.7244official
12llama-3.3-70b-instruct0.7191official
13llama-3.1-70b-instruct0.7183official
14qwen2.5-32b-instruct0.711official
15mistral-7b-instruct0.7038official
16olmo-7b-0724-instruct0.6598official
17llama-3.1-70b-base0.6398official
18llama-3.1-8b-base0.5653official
19deepseek-r1-distill-llama-70b0.525official
20s1.1-32b0.4755official

underspecified_context_f1

Measures whether the model asks for or abstains over missing context instead of inventing the omitted details.

RankModelValueRelative performanceProvenance
1qwen2.5-32b-instruct0.7403official
2gpt-4o0.7169official
3llama-3.1-405b-instruct0.7157official
4gemini-1.5-pro0.6963official
5llama-3.1-70b-instruct0.6927official
6o10.6907official
7mistral-7b-instruct0.6605official
8llama-3.3-70b-instruct0.6586official
9llama-3.1-8b-instruct0.656official
10llama-3.1-tulu-3-70b-dpo0.6557official
11llama-3.1-tulu-3-70b-ppo-rlvf0.6233official
12llama-3.1-tulu-3-70b-sft0.5215official
13deepseek-r1-distill-llama-70b0.5202official
14llama-3.1-70b-base0.5138official
15olmo-7b-0724-instruct0.5001official
16llama-3.1-8b-base0.4863official
17s1.1-32b0.4772official
18llama-3.1-tulu-3-8b-dpo0.4467official
19llama-3.1-tulu-3-8b-ppo-rlvf0.4319official
20llama-3.1-tulu-3-8b-sft0.3572official

underspecified_intent_f1

Measures whether the model asks for or abstains over unclear user intent instead of assuming a goal.

RankModelValueRelative performanceProvenance
1gpt-4o0.8152official
2llama-3.1-tulu-3-70b-dpo0.77official
3o10.7694official
4llama-3.1-tulu-3-70b-ppo-rlvf0.7587official
5llama-3.1-tulu-3-8b-ppo-rlvf0.7569official
6qwen2.5-32b-instruct0.7537official
7llama-3.1-tulu-3-70b-sft0.7492official
8llama-3.1-tulu-3-8b-dpo0.7425official
9gemini-1.5-pro0.7403official
10llama-3.1-8b-instruct0.7338official
11llama-3.1-tulu-3-8b-sft0.7311official
12mistral-7b-instruct0.725official
13llama-3.3-70b-instruct0.7242official
14llama-3.1-405b-instruct0.7024official
15llama-3.1-70b-instruct0.702official
16olmo-7b-0724-instruct0.6654official
17deepseek-r1-distill-llama-70b0.6276official
18s1.1-32b0.5964official
19llama-3.1-70b-base0.5544official
20llama-3.1-8b-base0.5494official