Model failure · Republic of Sanderia 2024 (B) · USD 2029 Notes · CDSBench #205 · benchmark v0.1
GPT answered Eligible; the verified outcome was Review
Question
Applying the standard Deliverable Obligation Characteristics (Not Subordinated, Specified Currency, Not Contingent, Transferable, Maximum Maturity 30 years, Not Bearer), is the instrument a Deliverable Obligation for this credit event?
- Verified answer
- Review Documents conflict on maturity; the conflict must be resolved before a conclusion.
- GPT answered
- Eligible All deterministic characteristics satisfied.
- Decisive issue
- Deliverability
- Evidence for the label
- GPT cited
- 1 non-existentevidence score 0%
Failure classification
Hallucinated evidence
- Hallucinated evidence. 1 cited section id(s) do not exist in the case record.
- Retrieval failure. None of the cited sections is among those that establish the answer; the relevant evidence was not retrieved.
- Failed to notice contradiction. Documents conflict on a material fact; the answer did not flag the conflict.
- Overconfident. Confidence 0.92 on an answer the record cannot support.
Classifications are produced by the failure analyzer against the verified label and the evidence map; the model never grades itself. Failure taxonomy →
Everyone else
The same question, every other participant
- MastermindCorrectanswered Review
- ClaudeCorrectanswered Review
- GeminiIncorrectanswered Eligibleits failure page →
Related failures
GPT also missed at least 6 other deliverability questions
- #57 Brightwater Energy Corp 2024Wrong final answer
- #67 Republic of Sanderia 2025Hallucinated evidence
- #76 Meridian Telecom S.A. 2026Hallucinated evidence
- #135 Solvane Airlines S.A. 2027Hallucinated evidence
- #195 Brightwater Energy Corp 2028 (B)Hallucinated evidence
- #214 Meridian Telecom S.A. 2025 (B)Hallucinated evidence
Benchmark version
v0.1
This page records GPT’s answer from its latest official run on the frozen benchmark v0.1. If the model is re-run and the answer changes, the page shows the new result and the model page keeps the full run history. Corrections policy →
Same kind of question on your own documents?
Mastermind runs the deterministic tests, the skeptic and the evidence verification on your record. Private documents never enter the benchmark.
Analyze your own documents →