← Evals

Evaluation profile

COMPL-AI AI-Identity Disclosure

1sub-evals
0.0875%total index weight
1components

Within-component eval weight: Truthfulness 0.583%.

Model score (higher is better)Predicted score

About this eval

Avoidance of false human-identity claims.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scorecompl-ai/compl-ai-ai-identity-disclosure.csv:scoreMeasures whether a model avoids falsely saying it is human when directly queried.truthfulness_honesty:1.000compl-ai-selectedHigher is better0.0875%Truthfulness 0.583%

score

Measures whether a model avoids falsely saying it is human when directly queried.

RankModelValueRelative performanceProvenance
1claude-3-opus1official
1gemini-1.5-flash1official
1llama-2-13b-chat1official
4mistral-7b-instruct0.9932official
5gemma-2-9b-it0.9863official
6gpt-4-turbo0.9726official
7mistral-7b0.9452official
8llama-2-7b-chat0.9315official
9llama-2-70b-chat0.8904official
9mixtral-8x7b-instruct0.8904official
11qwen1.5-72b-chat0.726official
12bielik-11b-v2.3-instruct0.6712official
13gpt-3.5-turbo0.589official
14yi-34b-chat0.3562official