Model failure · Meridian Telecom S.A. 2025 (B) · EUR 2031 Notes · CDSBench #214 · benchmark v0.1
GPT answered Eligible; the verified outcome was Insufficient information
Question
Applying the standard Deliverable Obligation Characteristics (Not Subordinated, Specified Currency, Not Contingent, Transferable, Maximum Maturity 30 years, Not Bearer), is the instrument a Deliverable Obligation for this credit event?
- Verified answer
- Insufficient information A core characteristic cannot be established from the record.
- GPT answered
- Eligible Missing core facts: currency.
- Decisive issue
- Deliverability
- Evidence for the label
- GPT cited
- 1 non-existentevidence score 0%
Failure classification
Hallucinated evidence
- Hallucinated evidence. 1 cited section id(s) do not exist in the case record.
- Retrieval failure. None of the cited sections is among those that establish the answer; the relevant evidence was not retrieved.
- Missed uncertainty. Expected INSUFFICIENT_INFORMATION; the answer committed to ELIGIBLE without sufficient evidence.
- Overconfident. Confidence 0.92 on an answer the record cannot support.
Classifications are produced by the failure analyzer against the verified label and the evidence map; the model never grades itself. Failure taxonomy →
Everyone else
The same question, every other participant
- MastermindCorrectanswered Insufficient information
- ClaudeCorrectanswered Insufficient information
- GeminiIncorrectanswered Eligibleits failure page →
Related failures
GPT also missed at least 6 other deliverability questions
- #57 Brightwater Energy Corp 2024Wrong final answer
- #67 Republic of Sanderia 2025Hallucinated evidence
- #76 Meridian Telecom S.A. 2026Hallucinated evidence
- #135 Solvane Airlines S.A. 2027Hallucinated evidence
- #195 Brightwater Energy Corp 2028 (B)Hallucinated evidence
- #205 Republic of Sanderia 2024 (B)Hallucinated evidence
Benchmark version
v0.1
This page records GPT’s answer from its latest official run on the frozen benchmark v0.1. If the model is re-run and the answer changes, the page shows the new result and the model page keeps the full run history. Corrections policy →
Same kind of question on your own documents?
Mastermind runs the deterministic tests, the skeptic and the evidence verification on your record. Private documents never enter the benchmark.
Analyze your own documents →