← Models

Model profile

Doubao 1.5 Thinking Pro

ByteDancedeveloper
2025-04-15release date
Not rankedSafety rank
#92 / 645Freedom rank

Evidence summary

Safety. Doubao 1.5 Thinking Pro does not meet the evidence gate for a Safety rank. Published Safety evidence spans 1 eval lineages and 4 of 7 components. Its strongest relative result is Concordia — Shutdown-Resistance (safety_score, #1 of 53); its weakest is Concordia — SciKnowEval-ChemicalHarmfulQA (safety_score, #63 of 63).

Freedom. Doubao 1.5 Thinking Pro has an estimated Freedom rank of #92; its 90% source-sensitivity interval is #16–#458. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is Concordia — SciKnowEval-ChemicalHarmfulQA (safety_score, #1 of 63); its weakest is Concordia — SOSBench-Bio (safety_score, #7 of 63).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Concordia — Agentic-Misalignmentsafety_score#26 / 5480Source ↗official
Concordia — AirBench-Deceptionsafety_score#58 / 6346.67Source ↗official
Concordia — AirBench-Manipulationsafety_score#55 / 5650Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#61 / 6346.4Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#58 / 6368.13Source ↗official
Concordia — MASKsafety_score#48 / 6249.16Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#61 / 637.407Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#63 / 630.6608Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53100Source ↗official
Concordia — SOSBench-Biosafety_score#57 / 6319Source ↗official
Concordia — SOSBench-Chemsafety_score#58 / 6340.4Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.