Skip to content
CDSBench
Release analysis · draft, not yet published as a report

GPT-5 on CDSBench v0.1

GPT is reliable on seniority (100%) but evidence remains its weakest category at 0%; its citations need checking before use (14.9% non-existent evidence).

Score
61.5
Previous generation
Change
Headline

What the measurements say

  • GPT-5 scores 61.5 on 296 verified questions, 2nd of 3 frontier models on CDSBench v0.1.
  • This is the first GPT generation on this benchmark version, so there is no predecessor to compare with.
  • Weakest category: evidence at 0% (0/30), against 100% for Claude.
  • It cited evidence that does not exist in the record in 14.9% of answers, and answered wrongly at ≥ 80% confidence in 18.9%.
  • On the 56 verified questions it missed, Mastermind (CDSBench's own system, reference only, 91.2) answered 42 correctly and abstained on 4.
Categories

Accuracy by category

CategoryAcc.nΔ
Evidence0%30
Restructuring0%2
Uncertainty69%32
Deliverability73%30
Entity match93%30
Transferability93%30
Fact extraction96%56
Currency100%28
Maturity100%26
Seniority100%30
Multi-document100%2
Hardest failures

Highest confidence first

The founder publishes this as a versioned report from the workbench release workflow; it then lives under /research with PDF, citation and share actions. Model page →