Evaluation profile
HalluVerse-M3 Hallucination Recognition
1sub-evals
0.262%total index weight
1components
Within-component eval weight: Truthfulness 1.75%.
Model score (higher is better)Predicted score
About this eval
Multilingual recognition of controlled hallucinated edits in QA and summarization.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| hallucination_recognition_accuracyhalluverse-m3/halluverse-m3.csv:hallucination_recognition_accuracyMeasures multilingual recognition of controlled hallucinations. | truthfulness_honesty:1.000halluverse-m3 | Higher is better | 0.262% | Truthfulness 1.75% |
hallucination_recognition_accuracy
Measures multilingual recognition of controlled hallucinations.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 80.3 | official | |
| 2 | gpt-4.1 | 78.66 | official | |
| 3 | claude-3.5-sonnet | 76.97 | official | |
| 4 | gemini-2.5-pro | 76.26 | official | |
| 5 | deepseek-v2.5 | 73.76 | official | |
| 6 | gpt-4o-mini | 73.39 | official | |
| 7 | llama-3.3-70b-instruct | 70.53 | official | |
| 8 | phi-4 | 70.38 | official | |
| 9 | palm-2 | 69.97 | official | |
| 10 | qwen2.5-72b-instruct | 69.81 | official | |
| 11 | qwen2.5-32b-instruct | 66.84 | official | |
| 12 | gemma-2-27b-it | 62.1 | official | |
| 13 | qwen2.5-7b | 59.4 | official | |
| 14 | mistral-7b | 56.06 | official |