Skip to content
CDSBench
Benchmark category · v0.1 · 30 verified questions

Evidence

Citing the sections of the record that actually support the answer, and no others.

CSV
Ranking

Accuracy on evidence questions

Compare →
#ModelAccuracyCorrect
1Claude Sonnet 5100%30/30
2GPT -50%0/30
3Gemini 2.5 Pro0%0/30
Reference · CDSBench’s own system · not ranked
Mastermind100%30/30
Common failures

How models go wrong here

Wrong fact60
Overconfident60
Retrieval failure60

Across every model’s latest official run. Failure taxonomy →

Hardest questions

Most models wrong

Latest

Most recently published questions

Cases

30 cases with evidence questions

and 18 more in the case explorer.

Following this category sends an email only when meaningful new evidence questions are published, a model’s evidence accuracy moves materially, or a research report on it is published. No newsletter.