Skip to content
CDSBench
OpenAI · GPT-5 · benchmark v0.1 · 2026-09-15

GPT

CDSBench score61.5on 296 verified questions

Not currently inside Mastermind. The promoted Mastermind version (v0) uses a different analyst model; this page is what tells us whether that should change.

Scores (CSV)

Every failure with a permanent page →GPT vs ClaudeGPT vs GeminiCompare with Mastermind →

Accuracy
81%
Evidence
21%
cited the right sections
Complex reasoning
79%
multi-document, deliverability
Uncertainty
33%
abstains only when it should
Hallucination
14.9%
cites non-existent evidence
False confidence
18.9%
wrong at ≥ 80% confidence
Strongest

Where it is reliable

  • Multi-document100%n=2
  • Seniority100%n=30
  • Maturity100%n=26
Weakest

Where it fails

  • Evidence0%n=30
  • Restructuring0%n=2
  • Uncertainty69%n=32
Failure profile

Accuracy by category and question level

CategoryAcc.n
Evidence0%30
Restructuring0%2
Uncertainty69%32
Deliverability73%30
Entity match93%30
Transferability93%30
Fact extraction96%56
Currency100%28
Maturity100%26
Seniority100%30
Multi-document100%2
LevelAcc.n
Uncertainty69%32
Deliverability73%30
Multi document25%8
Fact extraction63%86
Rule application100%112
Document interpretation100%28
Failure types

How it goes wrong

Evidence selection failure150
Overconfident56
Hallucinated evidence44
Retrieval failure36
Wrong fact30
Wrong final answer12
Unsupported claim10
Failed multi document reasoning8
Missed uncertainty6
Failed to notice contradiction2
Performance over time

All official runs on this benchmark version

5075100GPT-52026-09-152026-09-15
DateVersionOverallAccuracy
2026-09-15 17:16GPT-561.581%
2026-09-15 19:59GPT-561.581%
2026-09-15 20:14GPT-561.581%
2026-09-15 20:36GPT-561.581%