Values evaluation profile
Taiwan Sovereignty Benchmark Pro
About this evaluation
Mean Chinese/English prompt pass rates are shown as a compact headline because the prompts are paired translations scored by the same keyword rubric. Chinese and English rates, consistency, warning rates, soft-censorship match rates, API-error rates, and language gaps remain separately published. Repeated runs are not averaged: the all-error GPT-4o Mini run is excluded and retained in the audit artifact.
This is compatibility with an explicitly pro-Taiwan, source-authored keyword rubric—not geopolitical truth, neutrality, democracy quality, or generic model safety. Neutral quotation or discussion of PRC claims can trigger a fail; expected factual elements are not scored; soft-censorship phrases do not themselves fail; most cells are one sample from mutable hosted endpoints. Qwen3 Max API failures are preserved separately and its pass rate must not be read as pure narrative agreement.
Original source ↗ All values evaluations Download model results Download language results Download item results
Model results
| Model | Pro-Taiwan rubric compatibility | Warning-phrase rate | Soft-censorship rate | API-error rate | Configurations |
|---|---|---|---|---|---|
| Claude 3.5 Sonnet | 90 | 0 | 0 | 0 | 1 |
| Claude Opus 4.5 | 90 | 0 | 5 | 0 | 1 |
| Claude Sonnet 4.5 | 70 | 0 | 0 | 0 | 1 |
| Deepseek V3 | 10 | 40 | 0 | 0 | 1 |
| Gemini 2.0 Flash | 65 | 5 | 0 | 0 | 1 |
| Gemini 3 Pro Preview | 60 | 15 | 0 | 0 | 1 |
| GPT 4O | 90 | 0 | 0 | 0 | 1 |
| GPT 4O Mini | 100.0 | 0 | 0 | 0 | 1 |
| GPT 5.2 | 85 | 10 | 0 | 0 | 1 |
| Grok 3 | 55 | 20 | 10 | 0 | 1 |
| Kimi K2.5 | 15 | 40 | 0 | 0 | 1 |
| Llama 3.3 70B Instruct | 90 | 0 | 0 | 0 | 1 |
| Minimax M2 | 55 | 35 | 0 | 0 | 1 |
| Mistral Large 3 | 45 | 30 | 5 | 0 | 1 |
| Qwen2.5 72B Instruct | 20 | 10 | 5 | 0 | 1 |
| Qwen3 Max | 0 | 20 | 0 | 35 | 1 |
Dimensions
| Measure | Family | Native scale |
|---|---|---|
| Pro-Taiwan rubric compatibility | Taiwan sovereignty framing | 0 to 100 |
| Warning-phrase rate | Response behavior | 0 to 100 |
| Soft-censorship rate | Response behavior | 0 to 100 |
| API-error rate | Response behavior | 0 to 100 |