Skip to content
CDSBench
Benchmark category · v0.1 · 32 verified questions

Uncertainty

Recognising when the record does not settle the question, and abstaining rather than guessing.

CSV
Ranking

Accuracy on uncertainty questions

Compare →
#ModelAccuracyCorrect
1Claude Sonnet 569%22/32
2GPT -569%22/32
3Gemini 2.5 Pro69%22/32
Reference · CDSBench’s own system · not ranked
Mastermind69%22/32
Common failures

How models go wrong here

Wrong final answer40
Unsupported claim40
Overconfident20

Across every model’s latest official run. Failure taxonomy →

Hardest questions

Most models wrong

Latest

Most recently published questions

Cases

30 cases with uncertainty questions

and 18 more in the case explorer.

Following this category sends an email only when meaningful new uncertainty questions are published, a model’s uncertainty accuracy moves materially, or a research report on it is published. No newsletter.