Skip to content
CDSBench
Benchmark category · v0.1 · 30 verified questions

Deliverability

The full deliverable-obligation analysis: every characteristic tested together on one instrument and one credit event.

CSV
Ranking

Accuracy on deliverability questions

Compare →
#ModelAccuracyCorrect
1Claude Sonnet 583%25/30
2GPT -573%22/30
3Gemini 2.5 Pro73%22/30
Reference · CDSBench’s own system · not ranked
Mastermind83%25/30
Common failures

How models go wrong here

Hallucinated evidence68
Overconfident16
Retrieval failure14
Unnecessary review10
Evidence selection failure9
Missed uncertainty8

Across every model’s latest official run. Failure taxonomy →

Hardest questions

Most models wrong

Latest

Most recently published questions

Cases

30 cases with deliverability questions

and 18 more in the case explorer.

Following this category sends an email only when meaningful new deliverability questions are published, a model’s deliverability accuracy moves materially, or a research report on it is published. No newsletter.