Release analysis · draft, not yet published as a report
GPT-5 on CDSBench v0.1
GPT is reliable on seniority (100%) but evidence remains its weakest category at 0%; its citations need checking before use (14.9% non-existent evidence).
Score
61.5
Previous generation
—
Change
—
Headline
What the measurements say
- GPT-5 scores 61.5 on 296 verified questions, 2nd of 3 frontier models on CDSBench v0.1.
- This is the first GPT generation on this benchmark version, so there is no predecessor to compare with.
- Weakest category: evidence at 0% (0/30), against 100% for Claude.
- It cited evidence that does not exist in the record in 14.9% of answers, and answered wrongly at ≥ 80% confidence in 18.9%.
- On the 56 verified questions it missed, Mastermind (CDSBench's own system, reference only, 91.2) answered 42 correctly and abstained on 4.
Categories
Accuracy by category
| Category | Acc. | n | Δ |
|---|---|---|---|
| Evidence | 0% | 30 | — |
| Restructuring | 0% | 2 | — |
| Uncertainty | 69% | 32 | — |
| Deliverability | 73% | 30 | — |
| Entity match | 93% | 30 | — |
| Transferability | 93% | 30 | — |
| Fact extraction | 96% | 56 | — |
| Currency | 100% | 28 | — |
| Maturity | 100% | 26 | — |
| Seniority | 100% | 30 | — |
| Multi-document | 100% | 2 | — |
Hardest failures
Highest confidence first
- #9 Republic of Norland 2024 — answered DOC-1#S1, verified DOC-1#S4 · Retrieval failure
- #19 Halvard Steel AG 2025 — answered DOC-1#S1, verified DOC-1#S4 · Retrieval failure
- #29 Kingdom of Vestmark 2026 — answered DOC-1#S1, verified DOC-1#S4 · Retrieval failure
- #39 Orinoco Retail Holdings plc 2027 — answered DOC-1#S1, verified DOC-1#S4 · Retrieval failure
- #49 Federal Republic of Aurelia 2028 — answered DOC-1#S1, verified DOC-1#S4 · Retrieval failure
- #56 Brightwater Energy Corp 2024 — answered No, verified Yes · Failed multi document reasoning
The founder publishes this as a versioned report from the workbench release workflow; it then lives under /research with PDF, citation and share actions. Model page →