Skip to content
CDSBench
Release analysis · draft, not yet published as a report

Claude Sonnet 5 on CDSBench v0.1

Claude is reliable on entity match (100%) but transferability remains its weakest category at 57%; its citations need checking before use (6.1% non-existent evidence).

Score
91.2
Previous generation
Change
Headline

What the measurements say

  • Claude Sonnet 5 scores 91.2 on 296 verified questions, 1st of 3 frontier models on CDSBench v0.1.
  • This is the first Claude generation on this benchmark version, so there is no predecessor to compare with.
  • Weakest category: transferability at 57% (17/30).
  • It cited evidence that does not exist in the record in 6.1% of answers, and answered wrongly at ≥ 80% confidence in 0.0%.
  • On the 28 verified questions it missed, Mastermind (CDSBench's own system, reference only, 91.2) answered 0 correctly and abstained on 18.
Categories

Accuracy by category

CategoryAcc.nΔ
Transferability57%30
Uncertainty69%32
Deliverability83%30
Currency100%28
Evidence100%30
Maturity100%26
Seniority100%30
Fact extraction100%56
Entity match100%30
Multi-document100%2
Restructuring100%2
Hardest failures

Highest confidence first

The founder publishes this as a versioned report from the workbench release workflow; it then lives under /research with PDF, citation and share actions. Model page →