Skip to content
CDSBench
Anthropic · Claude Sonnet 5 · benchmark v0.1 · 2026-09-15

Claude

CDSBench score91.2on 296 verified questions

Used inside Mastermind. The promoted Mastermind version (v0) calls Claude Sonnet 5 as its analyst and skeptic model. If a later version of this model measures better here, a new Mastermind version can adopt it — through the same gates, and reported on the Mastermind page.

Scores (CSV)

Every failure with a permanent page →Claude vs GPTClaude vs GeminiCompare with Mastermind →

Accuracy
91%
Evidence
92%
cited the right sections
Complex reasoning
73%
multi-document, deliverability
Uncertainty
94%
abstains only when it should
Hallucination
6.1%
cites non-existent evidence
False confidence
0.0%
wrong at ≥ 80% confidence
Strongest

Where it is reliable

  • Restructuring100%n=2
  • Multi-document100%n=2
  • Entity match100%n=30
Weakest

Where it fails

  • Transferability57%n=30
  • Uncertainty69%n=32
  • Deliverability83%n=30
Failure profile

Accuracy by category and question level

CategoryAcc.n
Transferability57%30
Uncertainty69%32
Deliverability83%30
Currency100%28
Evidence100%30
Maturity100%26
Seniority100%30
Fact extraction100%56
Entity match100%30
Multi-document100%2
Restructuring100%2
LevelAcc.n
Uncertainty69%32
Deliverability83%30
Multi document75%8
Fact extraction100%86
Rule application100%112
Document interpretation61%28
Failure types

How it goes wrong

Hallucinated evidence18
Unnecessary review18
Wrong final answer10
Unsupported claim10
Performance over time

All official runs on this benchmark version

8090100Claude Sonnet 2026-09-152026-09-15
DateVersionOverallAccuracy
2026-09-15 17:16Claude Sonnet 591.291%
2026-09-15 19:59Claude Sonnet 591.291%
2026-09-15 20:14Claude Sonnet 591.291%
2026-09-15 20:36Claude Sonnet 591.291%